Claude Opus 5's crown collecting entered its third straight day — and its main victim is now its own stablemate
The Aggregate Digest — Monday, July 27, 2026
Claude Opus 5's crown collecting entered its third straight day — and its main victim is now its own stablemate. After sweeping GDPval, ARC-AGI-3 and the Vals AI suite on Friday and taking debate and longform-writing boards Saturday, Opus 5 today added Chatbot Arena (Document), edging Claude Opus 4.6 by ten Elo (1,520 vs 1,510), and pried two crowns straight from Claude Fable 5: Gert Labs Rankings (75.77 vs 73.38 GScore) and GBA Eval, where its 0.796 emulator score clears Fable's 0.745. It also slotted in at #2 on MineBench (2,043 Elo, behind GPT 5.6 Sol Pro's 2,086) and #5 on Design Arena (3D). The family scoreboard stays tight: Fable 5 holds #6 overall, Opus 5 sits #7 — and Fable answered by taking MWS Vision Bench at 79.9%, a tenth of a point ahead of GPT-5.6 Sol, dethroning Cotype Pro 3.
The day's quietest rout happened in telecom. OTel-2.0-LLM-31B-IT, a 31B domain specialist, took the GSMA Open-Telco aggregate (90.27%) and all six sub-boards — TeleTables, 3GPP, TeleMath, TeleQnA, TeleLogs and srsRAN-Bench — from its own 8.3B predecessor. On TeleTables the gap is brutal: 79.76% vs 61.8%, with the best generalist, gemini-3.1-pro-preview, back at 48%.
At the other end of the table, Qwen3 Coder Next — #229 overall — sniped three Vector Eval boards by hairline margins: MATH from o3-mini (+0.91), InterCode-CTF from Claude 3.5 Sonnet (+0.33) and GPQA-D from o1 (+0.12). And on PutnamBench, Humanfia (w/ GPT 5.5) squeezed past Aleph Prover 670 problems to 668 in the theorem-proving arms race.
Eight benchmarks joined the corpus, and the pairs are the story. On CodeRouterBench, Claude Opus 4.6 leads the in-distribution split (43.83%), but shift to unseen repositories and GPT-5.4 takes over (64.2% vs 63.64) — a routing study where the winner flips with the distribution. WorldTravel-Bench is humbling: GPT-5.2, #42 overall, leads both settings, yet only 32.67% of its text-mode itineraries are feasible (19.33% multimodal) against a 77.9% human reference. Humanlaya's agent pair rounds it out: Claude Opus 4.7 tops Biomni's biomedical data-analysis tasks (73.34), GPT-5.4 the Excel board (72.92%).
Elsewhere. Kimi K3 (#11 overall) posted #2 on SWE-Milestone (46.75%, behind Claude Opus 4.8's 51.84) and #4 on MWS Vision Bench, and GDPevo arrived measuring agent self-evolution on 240 business tasks with Claude Opus 4.6 on top.