Claude Opus 5 took a fourth straight day of crowns, and its best one was for refusing to answer
The Aggregate Digest — Friday, July 31, 2026
Claude Opus 5 took a fourth straight day of crowns, and its best one was for refusing to answer. MathArena's ARXIV_FALSE June board hands models plausible but deliberately false proof statements lifted from June arXiv papers and scores whether they refuse and name the claim as false instead of fabricating a proof. Opus 5 posted 90.74 there, against 69.44 for the previous leader GPT-5.5 — a 21.3-point jump on a board of 16, and the largest single-board gain in today's digest. It is the least glamorous kind of win and probably the most useful one: the failure mode being measured is confident invention.
The same model added two more. ProfBench, at 99 models the broadest board it touched today, went to Opus 5 at 66.2, up from Claude Opus 4.7's 61.3 — an in-house handover, the #4 model overall taking a crown off the #36. On OckBench it reached 95.0 at extra-high effort, past Claude Opus 4.8's 91.5, with its own high and max settings filling the two places behind it.
It did not sweep. GPT-5.6 Sol, #8 overall, took the Vals AI Time Horizon Index at 13.0 from Claude Opus 4.8's 11.83, and won the day's most interesting new board outright.
That board is MathArena's ARXIVLEAN June: the same June arXiv problems as the prose track, but formalized in Lean, so a proof either machine-checks or it does not. GPT-5.6 Sol leads at 37.5 running via Codex, Opus 5 second at 31.25. Set that against the 90.74 on the prose sibling and the gap between recognizing a bad claim and constructing a verified good one is the whole story.
Three more new boards came from HumanCLAW-Bench, which drives a simulated humanoid from a frozen off-the-shelf VLM and gates each stage on the last. Gemini-3.1 leads all three, and the cascade is brutal: 64.9 for getting the target into view, 42.5 for walking to it and stopping within 20 cm, 16.8 for the full find-reach-sit sequence. No model in the field clears 20 on the end-to-end task.
Elsewhere, Kimi K3 placed 7th of 55 on Epoch AI's Apex Agents at 39.3, and edged its own default configuration on Tau3 Banking — 34.02 at low effort against 33.4 — by six tenths of a point.