Claude Opus 5 took five more crowns today, a fourth straight day of collecting — and three of the five came off its own stablemate
The Aggregate Digest — Tuesday, July 28, 2026
Claude Opus 5 took five more crowns today, a fourth straight day of collecting — and three of the five came off its own stablemate. The #6 model won FoodTruckBench outright, finishing the 30-day business simulation with $75,264 in net worth against GPT-5.5's $61,408, and took Chatbot Arena (Code) in its Max configuration at 1,725 Elo, 43 clear of Kimi K3. The other three were Claude Fable 5's: KernelBench Hub - CUDA, where Opus 5's best run reaches 103.67% of hardware roofline to Fable's 72.66; KernelBench Hub - Mega, 24.29× over the reference megakernel against 19.12×; and HiL-Bench, decided by two-thirds of a point (57.0 to 56.33) on whether an agent knows when to ask a human for help.
Fable 5 did not surrender the whole GPU suite. It still holds KernelBench Hub - Hard, 55.21% of roofline to Opus 5's 52.61 — the one board in the family fight that stayed put. And Opus 5's day was not all crowns: it arrived on LisanBench at #38 of 154, on TaxCalcBench at #8 of 14 with 18.0% of returns strictly correct, and third of 26 at heads-up poker.
The better fight was on VANTAGE-Bench, where a Flash-tier model outran the frontier. Gemini 3.6 Flash — #35 overall, thirty-odd places below its opposition — took the Overall board at 69.52%, beating #10-ranked GPT-5.6 Sol by less than half a point, and swept Event Verification (82.0%), Semantic (79.41%) and Single Object Tracking (75.99%) with it. Sol answered on Video QA (77.9%) and Temporal (46.25%), and neither of them got Spatial: Qwen3.5-27B holds it at 79.04%. Temporal is the slice that humbles all seventeen entrants — the best score there is 46.25% while every other slice tops 75%.
In computer use, Intelligence-Indeed Agent posted 90.19% on OSWorld's 369 desktop tasks, 6.55 points ahead of Pointer Agent w/ Opus 4.7 and the first score on that 64-model board to clear 90.
One benchmark joined the corpus, and it is a hard one: SpreadsheetBench v2, where Claude Opus 4.6 leads eight models at 34.89% and no one else clears 27%.
Elsewhere. Kimi K3 (#11 overall) landed #8 on Mercor APEX (39.3% best pass@1 of 52 models) and #6 on the offline TrackingAI IQ test, and GPT-5.6 Sol squeezed into the top ten on Kaggle FACTS at 52.41%.