The Aggregate Digest logo

The Aggregate Digest

Archives
Log in
Subscribe
September 3, 2026

Claude Fable 5.1 is now the best measured model as well, taking the crown Claude Mythos 5 had held for 86 days, 1814 to 1807

The Aggregate Digest — Thursday, September 3, 2026

Claude Fable 5.1 is now the best measured model as well, taking the crown Claude Mythos 5 had held for 86 days, 1814 to 1807. Claude Mythos 5 is private, so since June 9 the strongest model on the site was one you could not use. The strongest model we have measured is now the same model that leads the best available model table. Yesterday the gap was two points, inside the five-point margin a crown needs to move, and today it is seven.

The extra points came from 24 new results. Claude Fable 5.1 is first of 124 on Chatbot Arena (Code). On Epoch AI's ProofBench, where models write Lean 4 proofs of graduate-level theorems and a verifier checks them, it scores 100 percent and is first of 28. Claude Opus 5 has 99. It ties Claude Opus 5 at the top of Creative Writing (Lechmazur). On Epoch AI's CritPt it is second of 170, behind GPT-5.6 Sol. Kimi K3 stays best open-weights model at 1752.

Gemini 3.8 Flash arrived today with 82 results and is tied for third among available models with GPT-5.5 Pro at 1781, behind only Claude Fable 5.1 and Claude Opus 5. It is first of 604 on Artificial Analysis' GPQA Diamond, 198 graduate-level science questions on which PhD experts score 65 percent, at 95.25 percent. Grok 4.6 is second at 94.95. It is first of 252 on Artificial Analysis' MMMU-Pro, first of 43 on LVBench, which asks questions about videos averaging 68 minutes, and first of 55 on Vals AI Finance Agent v2, analyst work over SEC filings. LLM Stats' run of Terminal-Bench 2.1 has it first of 35, and Vals AI's run has it fourth of 60. It is 24th of 88 on Vals AI SWE-bench Verified and 15th of 595 on Artificial Analysis' Humanity's Last Exam.

Elsewhere. Muse Spark 1.3 has 30 results, five short of a ranking, and is first of 196 on Tau3 Banking, a set of banking customer-service tasks, ahead of Qwen 3.8 Max. Hy4 preview, on the site since Monday, is 18th among available models at 1733.

Read the full digest on The Aggregate →

Don't miss what's next. Subscribe to The Aggregate Digest:
Older → Claude Fable 5.1 arrived today and went straight to the top of the best available model table, ending Claude Opus 5's twelve-day hold at 1809 to 1786
aibenchmarks.dev
Twitter
Telegram
Powered by Buttondown, the easiest way to start and grow your newsletter.