The Aggregate Digest logo

The Aggregate Digest

Archives
Log in
Subscribe
July 25, 2026

Claude Opus 5 arrived and took every single lead change in today's digest — thirty crowns, one debutant

The Aggregate Digest — Saturday, July 25, 2026

Claude Opus 5 arrived and took every single lead change in today's digest — thirty crowns, one debutant. The new Anthropic model enters the overall ranking at #7 of 1,503 with an ELO of 2882, slotting in just below its stablemate Claude Fable 5 and just above GPT-5.4 Pro — and then spent its first day dethroning almost everyone above and around it.

The loudest result is ARC-AGI-3: Claude Opus 5 (High) posted 30.16% accuracy on a board whose previous best was GPT-5.6 Sol (Max) at 7.78% — nearly quadrupling the frontier on one of the corpus's hardest reasoning tests. On Vals AI SWE-bench Verified it became the first model to hit 97.0% resolved, nudging past GPT-5.6 Sol's 96.2%. The Vals AI suite fell almost wholesale: CorpFin v2, Finance Agent v2, MortgageTax, MedCode, MedScribe, ProofBench, Legal Research Bench, Code Migration, IOI, ProgramBench, MMLU-Pro, MMMU and SWE-bench Verified all now lead with Opus 5. On Vellum's Humanity's Last Exam it squeezed out Claude Mythos 5 by two tenths of a point, 64.7 vs 64.5, and on BenchLM it took the top slot at 85.9 ahead of Mythos 5. The Artificial Analysis Intelligence Index crown went to its Max Effort variant at 60.69.

The honest caveats: many of those margins are razor-thin — +0.09 on Vals AI MMLU-Pro, +0.57 on MMMU — and for all the crowns, the aggregate still ranks Opus 5 one notch below Claude Fable 5 overall. Its victims tell the story of the season: mostly Fable 5 and GPT-5.6 Sol, with Gemini 3.5/3.6 Flash, Grok 4.5, Muse Spark 1.1 and Claude Opus 4.7 losing a board each.

Four benchmarks joined the corpus today. LLM Stats added DeepSWE 1.1 (software engineering agents in the standardized mini-swe-agent harness; GPT-5.6 Sol leads at 73.0% across 18 models) and FrontierCode 1.1 (are agent-produced code changes actually mergeable, judged by unit tests and maintainer rubrics; Claude Fable 5 holds that one at 53.5% — the one new coding board Opus 5 didn't take). DecBench scores decompilation exactness, with GPT-5.6 Sol on top of a still-tiny two-model board. And Bongard Problems (Classic) — the human-designed Foundalis puzzle set where models must induce the visual rule separating six left from six right figures — opens as an all-Gemini affair: Gemini 3.1 Pro (Preview) leads at 55.4%, and no frontier model from any other lab has cracked it yet.

Read the full digest on The Aggregate →

Don't miss what's next. Subscribe to The Aggregate Digest:
← Newer Day two of Claude Opus 5, and the debutant is still collecting — two more crowns today, both lifted from its own stablemate, Claude Fable 5 Older → UniWorld-View walked onto the WorldScore board and swept it
aibenchmarks.dev
Twitter
Telegram
Powered by Buttondown, the easiest way to start and grow your newsletter.