The Aggregate Digest logo

The Aggregate Digest

Archives
Log in
Subscribe
August 21, 2026

Claude Opus 5 is now the best available model, after winning two agentic boards in this morning's Artificial Analysis refresh

The Aggregate Digest — Friday, August 21, 2026

Claude Opus 5 is now the best available model, after winning two agentic boards in this morning's Artificial Analysis refresh.

AA GDPval hands a model 220 real professional tasks drawn from 44 occupations, with shell and web access to do them. Claude Opus 5 came first there of 204 entries, taking the board from Qwen3.8 2.4T A95B. AA-Briefcase asks for agentic knowledge work over a briefcase of real files, scored by rubric. Claude Opus 5 passed 57.98% of those checks. Kimi K3, the previous leader, passed 50.97%. Artificial Analysis' own Intelligence Index also has Claude Opus 5 on top, across 597 entries.

On our table Claude Opus 5 now sits at 1792 across 231 benchmarks, which ends a six-day reign. Gemini 3.7 Flash had held best available model since August 15. GPT-5.5 Pro is second today at 1788. Gemini 3.7 Flash sits at 1784, tied for third with Claude Fable 5. Best measured model stays Claude Mythos 5, and best open-weights model stays Kimi K3.

The refresh was not all good news for Claude Opus 5. AA-Omniscience scores factual recall and hallucination over six thousand questions in six subjects, and Claude Fable 5 beat Claude Opus 5 in all six of them. Claude Opus 5 came second in business, health, humanities and science, third in software engineering, eighth in law.

Read the full digest on The Aggregate →

Don't miss what's next. Subscribe to The Aggregate Digest:
Older → GLM-5.3 enters the best-available ranking tied for 11th of 681, and its strongest results are on agent benchmarks
aibenchmarks.dev
Twitter
Telegram
Powered by Buttondown, the easiest way to start and grow your newsletter.