
Cohere Embed 5: Pro and Fast share one retrieval index
Cohere's new embedding family supports multimodal retrieval with Pro and Fast tiers that share an embedding space.
News & Insights
Reporting, explainers, and implementation-focused analysis on AI systems, automation, voice agents, and business delivery.

Google has announced Gemini 4 Argon with initial access for trusted cybersecurity defenders. Wider paid API and Google AI Ultra availability is coming later.
Model filters apply only to articles with verified model details.

Cohere's new embedding family supports multimodal retrieval with Pro and Fast tiers that share an embedding space.

GPT-6.1 Sol joins OpenAI's GPT-6 range. Here are the verified API details and the migration choices teams should check.

Sonnet 5.5 pricing and migration checks, plus a practical evaluation plan for everyday coding and document workflows.
Google's Live Avatar feature combines Gemini 3.8 Live with streaming video for enterprise conversations. Custom avatar creation requires allowlisting.

Google has released two speech generation models: Flash TTS for creative voice direction and Flash-Lite TTS for high-volume audio workflows.

FLUX 3 Action is an open-weight world action model for joint video and robot-action prediction. It needs a robotics evaluation, not an image-generation comparison.

GPT-6 Sol joins OpenAI's GPT-6 range. Here are the verified API details and the migration choices teams should check.

GPT-6 Luna joins OpenAI's GPT-6 range. Here are the verified API details and the migration choices teams should check.

Verified Opus 5.5 pricing and migration considerations, with a practical approach to evaluating cost per accepted result.

Xiaomi's dated announcement confirms the MiMo-V2.6 family, open checkpoints and a faster Pro serving option with distinct pricing.

SpaceXAI released Grok 4.7 for longer coding and knowledge tasks, with standard API pricing and a faster variant to compare on real workloads.

Qwen released an image generation and editing model with native transparency, local edits and support for up to ten reference images.

Qwen's interpretation release adds speaker attribution and bilingual output. Its language coverage differs between text and spoken translations.

Qwen's official release listing confirms Omni-Flash, focused on turning multimodal understanding into tool-assisted productivity work.

Google's two new live dialogue models pair spoken conversation with background tool execution. Developer rollout and enterprise preview access differ.

GPT-Live-1 separates a conversational voice layer from backend reasoning, with voice sessions priced at $0.05 per minute.

DeepSeek has launched DeepSeek-V4.1-Flash, a 552B-parameter mixture-of-experts model featuring a novel Causal Encoder–Decoder architecture with just 8B active input and 16B active output parameters. The release brings native vision, compresses KV cache storage by up to 8x, and slashes off-peak API pricing to $0.15 per million input tokens.

ChatGPT Images 2.5 brings two API models for generation and editing: Flare for everyday use and Sunburst for precision work.

OpenAI's life sciences model leaves research preview for approved organisations, with restrictions on internal research use.

AIMI recorded four distinct model routes arriving on OpenRouter during 4 September 2026. GPT-6 Astra led the late-evening wave, while Grok 4.3, Ling 3.0 Flash Sante and NVIDIA Nemotron 3.5 Content Safety appeared earlier.

Google's new weather model uses live satellite observations for hourly forecasts, with different spatial resolutions for surface and atmospheric variables.

OpenAI has officially launched GPT-6 Astra, its frontier flagship model featuring a 1,050,000-token context window, 128,000 max output tokens, and native tool-use for autonomous agent workflows.

Alibaba's Qwen team has launched Qwen3.8 Max, a 2.4T parameter Mixture-of-Experts model offering 1M token context, native video perception, and deep agent tool orchestration.

InclusionAI's official model card dates Ling 3.0 Flash Fin to 3 September 2026. OpenRouter then listed paid and zero-price routes for the 124-billion-parameter finance model.

Google's specialised cybersecurity model focuses on vulnerability discovery and patching. Access is through the Fairwind Program, rather than a general public rollout.

Meta Superintelligence Labs has published Muse Spark 1.3, an open-weights frontier agent model with 1,048,576 context tokens and comprehensive multimodal comprehension.

Google has released Gemini 3.8 Flash as a generally available model for long-horizon coding, autonomous agents and enterprise workflows. The API keeps the introductory price of Gemini 3.7 Flash, with a one-million-token input limit and 65,536-token output limit.

Anthropic announced Mythos 5.1 alongside Fable 5.1 in September. Its trusted-access restrictions matter: this is not a general public API launch.

Anthropic released Claude Fable 5.1 on 1 September 2026. Earlier that afternoon, OpenRouter added a route for Inception's Mercury 2.5 Preview. The timestamps are close, but the evidence does not show a response or shared launch.

OpenRouter added a paid route for IBM Granite 4.2 8B on 31 August. The hosted endpoint exposes a 131,072-token context window, reasoning and tool calling, six days after IBM released the open model family.

GLM-5.3-Flash is now listed in ClinePass after first appearing in OpenCode Go. The route additions expand access to Z.ai's native multimodal coding model, but they are provider availability events, not new maker releases.

Tencent has released and open-sourced Hy4 preview, a 770-billion-parameter mixture-of-experts model with 49 billion active parameters and a context window above one million tokens. AIMI saw it reach OpenRouter and OpenCode Go later the same day.

Google's developer video model adds scene continuation, first and last frame controls, low-resolution previews and 4K upscaling.

Google's dedicated speech-to-text model has separate routes for live speech and recorded audio. Its release predates the September Gemini audio announcements.

OpenRouter has added Qwen3.8 Flash, the production model Qwen says is served through QwenCloud with a one-million-token context and built-in tools. Qwen published the related Flash-Next open-weight preview on the same day.

AIMI saw GLM-5.3 Flash reach OpenCode Go, OpenRouter, Cloudflare Workers AI and Ollama Cloud on 26 August. The route wave expands access to Z.ai's model, but Z.ai has not announced a separate Flash launch.

DeepSeek has released V4 Flash Vision Exp, an experimental API model that adds image understanding to the V4 Flash family while keeping its text capabilities.

OpenCode Go now lists Ox Alpha Free as a limited-time model. The live route changed from x-preview-f-free to ox-alpha-free during the same monitoring window, while the model remains a stealth release with no public developer or model card.

Qwen has released Qwen3.8-27B under Apache 2.0. The dense vision-language model has a native 262,144-token context, adjustable reasoning, image and video input, and a live OpenRouter route while Qwen Cloud hosting remains pending.

Dots Studio has released dots3-note preview, its first open-weight dots3 model. The multimodal MoE has 280 billion total parameters, 16 billion active parameters, a 512K context window, and an Apache 2.0 release.

DeepSeek has released the GA version of V4 Pro for its app, web service and API. The 0813 model adds adjustable reasoning, native Responses API support, and a later hosted route on NVIDIA NIM.

Google has released Gemini 3.7 Flash as a stable model for coding, agent workflows and multimodal reasoning. It has a 1,048,576-token input limit, 65,536-token output, and direct API pricing from $0.75 per million input tokens through 2026.

xAI announced Grok 4.6 on 12 August 2026, tuned for long-running agents and ambitious visual work. It is live on OpenRouter with a 500,000-token context at $2 per million input tokens, and the release date is confirmed by the official announcement.

ByteDance released Seed 2.1 Turbo on 23 June 2026. Its OpenRouter route reached AZ Labs monitoring on 12 August with text, image and video input, a 262,144-token context, and pricing from $0.50 per million input tokens.

Meta Superintelligence Labs has released Muse Glimmer 30B, an Apache 2.0 open-weight model distilled from Muse Spark and tuned for local agent workflows on consumer GPUs. It covers tool use, long-context reasoning and multimodal input in a package that fits under 20 GB with 4-bit quantisation.

NVIDIA removed DeepSeek V4 Flash, DeepSeek V4 Pro and Mistral Medium 3.5 128B from the NIM serverless API on 7 August 2026. We confirmed the removals against the live endpoint and tracked where these models still run.

OpenRouter now lists InclusionAI's Ling 3.0 Tiny as a zero-price route with a 262K-token context window, switchable thinking and instant modes, and a 32K maximum output.

A practical model-routing experiment puts InclusionAI's Ling 3.0 Flash inside Grok CLI, with the request verified through OpenRouter and the same route working in Pi.

Anthropic has unveiled Claude Opus 5, its flagship frontier model delivering unprecedented coding capabilities, 1M context comprehension, and state-of-the-art agentic reasoning.

Google launched Gemini 3.6 Flash on 21 July 2026 alongside 3.5 Flash-Lite and 3.5 Flash Cyber, cutting output token usage by 17% against 3.5 Flash at $0.75 per million input tokens.

Moonshot AI has released Kimi K3, a 2.8-trillion parameter MoE powerhouse featuring 1M context, native video understanding, and leading mathematical reasoning.

A practical guide to running AI voice agents in South Africa while staying compliant with POPIA. Covers data residency, consent, storage, and how to design voice workflows that respect privacy obligations from day one.

Anthropic has introduced Claude Sonnet 5, offering 1M context processing, adaptive thinking effort controls, and unmatched throughput for software engineering.

OpenAI has unveiled GPT-5, its most advanced language model to date, featuring breakthrough reasoning capabilities that bring AI closer to human-level problem solving.

Google DeepMind launched Gemini 2.5 Pro with native multimodal reasoning, setting new benchmarks across coding, math, and scientific analysis tasks.

From fintech to healthcare, South African companies are rapidly integrating AI into their operations. We explore the trends driving adoption across the continent.

Introducing Aura, our AI-powered booking assistant that handles scheduling, reminders, and client communications autonomously for service businesses.

Learn how to design, build, and deploy AI voice agents that deliver natural conversations and measurable business results.

Anthropic's latest model, Claude 4, pushes the frontier of responsible AI development with state-of-the-art performance on safety and helpfulness benchmarks.

As AI adoption accelerates across Africa, governments are developing regulatory frameworks. We examine the policies and what they mean for businesses.

From contract analysis to legal research, AI is reshaping how law firms operate. Discover the practical applications driving efficiency in the legal sector.

We're bringing our AI expertise to the education sector, partnering with institutions to build intelligent tutoring systems and administrative automation.