This week in AI models: DeepSeek-V4.1-Flash, GPT Image 2.5, MAI-Image-2.6 +7 more
What shipped
10 new releases on the tracker — what actually changed versus the model each one replaces, in the makers' own numbers.
DeepSeek-V4.1-Flash
DeepSeek · DeepSeek Flash line · lightweight · open weights · Sep 10, 2026
A new encoder-decoder design that activates only 8B parameters on input and 16B on output, beating V4-Flash on every agentic test DeepSeek published while keeping the MIT licence and 1M-token context — though the API price more than doubles.
Compared with DeepSeek-V4-Flash-0731 · the model it replaces
| Finishes 9 in 10 multi-step tasks at the command line Terminal-Bench 2.1 — running real commands to complete a task | 90.6% ▲ up from 82.7% · +8 points |
| Five points smarter on the all-round score Artificial Analysis Intelligence Index — a broad average across many tests | 40 ▲ up from 35 · +5 points |
| Costs developers more than twice as much to run Price per 1M tokens — what developers pay, input / output, at peak hours | $0.30 / $1.20 price rose — was $0.14 / $0.28 |
Sources: DeepSeek · Hugging Face · Artificial Analysis
GPT Image 2.5
OpenAI · GPT Image line · image model · Sep 8, 2026
Ships as two variants in ChatGPT and the API: Flare, with higher quality than GPT Image 2 at 50% lower latency, and Sunburst, a slower model for precise editing work.
Compared with GPT Image 2 · the model it replaces
| Ranks first in blind image comparisons Artificial Analysis Image Arena — Elo from people comparing images blind | 1187 ▲ Flare variant · up from 1171 · +16 Elo |
Sources: OpenAI · OpenAI API (Flare) · Wikipedia · Artificial Analysis
MAI-Image-2.6
Microsoft · MAI Image line · image model · Sep 4, 2026
No. 2 on the Arena text-to-image leaderboard, +79 Elo over MAI-Image-2.5 across all eight categories.
Compared with MAI-Image-2.5 · the model it replaces
| Ranks fourth in blind image comparisons Artificial Analysis Image Arena — Elo from people comparing images blind | 1144 ▲ up from 1105 · +39 Elo |
Sources: Microsoft AI · Artificial Analysis
GPT-6 Astra
OpenAI · GPT line · flagship · Sep 3, 2026
OpenAI's new flagship: far fewer made-up answers and a third of the tokens on coding-agent work, but the same overall score as GPT-5.6 Sol — at 2.5 times the price.
Compared with GPT-5.6 Sol · the model it replaces
| Invents facts on half as many trick questions AA hallucination rate — how often it makes up an answer instead of admitting it doesn't know | 51% ▲ down from 92% · 41 points fewer |
| Uses about a third of the tokens on coding-agent tasks AA Coding Agent Index, max effort — text consumed per task | ~⅓ tokens ▲ vs GPT-5.6 Sol · 70% more token-efficient |
| Costs 2.5 times more to run than Sol Price per 1M tokens — what developers pay, input / output | $10 / $50 up from $4 / $20 · 2.5× the price |
Sources: OpenAI · Artificial Analysis · CNBC
Gemini 3.8 Flash
Google · Gemini Flash line · lightweight · Sep 2, 2026
Google's new Flash: a big jump on command-line tasks and a few points smarter, at the same price — though it burns about 30% more tokens per task, and the price doubles in 2027.
Compared with Gemini 3.7 Flash · the model it replaces
| Finishes 9 in 10 multi-step tasks at the command line Terminal-Bench 2.1 — running real commands to complete a task | 90.8% ▲ up from 81.6% · +9 points |
| A few points smarter on the all-round score Artificial Analysis Intelligence Index — a broad average across many tests | 59 ▲ up from 56 · +3 points |
| Same price as 3.7 Flash — until the end of 2026 Price per 1M tokens — what developers pay, input / output | $0.75 / $3.75 unchanged · doubles to $1.50 / $7.50 on Jan 1, 2027 |
Sources: Google · Vellum · 9to5Google
Muse Spark 1.3
Meta · Muse line · flagship · Sep 2, 2026
Meta's flagship edges closer to the leaders: four points up on the all-round score and clear gains on office and terminal tasks, at the same price and the same 1M-token context.
Compared with Muse Spark 1.2 · the model it replaces
| Four points higher on the all-round intelligence score Artificial Analysis Intelligence Index — a broad average across many tests | 61 ▲ up from 57 · +4 points |
| Does better on realistic office work GDPval-AA v2 — real knowledge-work tasks, scored as an Elo rating | 1709 ▲ up from 1615 · +94 Elo |
| Same price as Muse Spark 1.2 Price per 1M tokens — what developers pay, input / output | $1.25 / $4.25 unchanged from Muse Spark 1.2 |
Sources: Meta AI Research · Artificial Analysis · Axios
Claude Fable 5.1
Anthropic · Claude Fable line · flagship · Sep 1, 2026
Anthropic's top model, refreshed: big jumps on long coding and scientific-research tasks, and cached input now costs 75% less — same headline price as Fable 5.
Compared with Claude Fable 5 · the model it replaces
| Finishes over half of long multi-step coding tasks Terminal-Bench 4.0 — long agentic coding sessions in a terminal | 55.8% ▲ up from 42.0% · +14 points |
| Doubles its score on agentic scientific research Terminal-Bench-Science 0.1 — running real research workflows end to end | 52.6% ▲ up from 24.7% · more than doubled |
| Same price — and cached input now costs 75% less Price per 1M tokens — what developers pay, input / output | $10 / $50 same as Fable 5 · cache reads now $0.25 |
Sources: Anthropic · VentureBeat
Gemini Omni 1.1 Flash
Google · Gemini Omni line · video model · Aug 27, 2026
Scene extension now reads 10 seconds of prior footage (up from 1) for clips up to 40 seconds, adds first-and-last-frame camera control, 4K upscaling and a 360p draft mode that renders up to 60% faster at a third of the cost.
Sources: Google · Google DeepMind
Qwen3.8-Flash-Next
Alibaba · Qwen Flash line · lightweight · open weights · Aug 26, 2026
Open-weight mixture-of-experts model with 125B total parameters but only 6B active per token, a 262K native context window and a built-in vision encoder, previewing the Qwen4 architecture.
| Scores 40 on the all-round score Artificial Analysis Intelligence Index — a broad average across many tests | 40 ✦ first model in this line — nothing earlier to compare |
Sources: Hugging Face · GitHub (QwenLM) · TechNode · Artificial Analysis
Wan 3.0
Alibaba · Wan line · video model · open weights · Aug 24, 2026
Doubles the maximum clip length to 30 seconds in a single pass, accepts documents (PDF, PPT, DOC) as input alongside text, image, audio and video, and improves face realism and reference consistency.
Compared with Wan 2.7 · the model it replaces
| Generates clips twice as long in one go Maximum clip length — seconds of video from a single generation | 30 seconds ▲ up from 15 seconds · 2× |
| Costs from 5 cents per second of video Price per second of video — 480p / 720p / 1080p | $0.05 / $0.10 / $0.20 ✦ launch list price by resolution |
Sources: Alibaba Cloud · Alizila · GitHub (Alibaba Cloud)
Coming next
Where the flagship lines stand in their own release rhythm. No dates — nobody can promise those — just the facts.
| Gemini · Google · last: Gemini 3.1 Pro | DUE 207 days since the last release — usual gap is 132 |
| Mistral Large · Mistral · last: Mistral Large 3 | DUE 286 days since the last release — usual gap is 149 |
| Claude Opus · Anthropic · last: Claude Opus 5 | MID-CYCLE 52 of ~70 days into its usual cycle |
| Muse · Meta · last: Muse Spark 1.3 | MID-CYCLE 12 of ~28 days into its usual cycle |
| Grok · xAI · last: Grok 4.6 | MID-CYCLE 33 of ~72 days into its usual cycle |
Every line we track, where it stands and what its last release changed: Next AI Model.
Updated automatically every morning · every number links to its source · one email a week, only when something shipped. No filler.