Hand a model image tools and its score on the same exam multiplies eight-and-a-half-fold
The Aggregate Digest — Thursday, July 23, 2026
Hand a model image tools and its score on the same exam multiplies eight-and-a-half-fold. That is the finding of the day's most interesting debut, a paired benchmark called ActiveVision — an "exam for active observers" that asks models to scan, traverse and transform images across 85 items humans solve at 96%. As pure chain-of-thought observers they flunk it: GPT-5.5 leads at 10.6%, Gemini 3.5 Flash takes 8.2%, Claude Fable 5 manages 5.9%. The paired agents board hands the same 85 items to tool-using harnesses, and Fable 5 jumps from 5.9 to 50.6% — with GPT-5.5 Codex at 37.6 and Claude Opus 4.8 at 24.7. Even the multiplied scores sit at barely half the human mark.
That crown headlines a strong day for Claude Fable 5 (#6 overall). It opened another debut board on top — KernelBench Hub's CUDA suite, which scores agentic kernel optimization as a percent of hardware roofline, at 72.66, thirty points clear of Kimi K3 (256k)'s 42.46 — and took Vals AI Public Benefits Bench off Kimi K3, 70.43 to 68.27. The counterweight: 24.0% pass@1 and fifth place on CHI-Bench, a board its stablemate Claude Opus 4.8 leads at 37.3.
Gemini 3.6 Flash, which debuted yesterday at #23 with a single crown, is already #17 overall and collected four more: Kaggle FACTS Multimodal at 50.24 (off Gemini 2.5 Pro's 46.86), AutomationBench in its High configuration at 19.8% pass (off GPT-5.6 Sol (Max), 18.1), and both 200-plus-model AI for Education boards — Pedagogy Maths at 95.24 and SEND at 88.53.
Kimi K3 (#11) posted six Agent Arena rows that cut both ways: third of 38 on Confirmed Success (14.0%) and on Praise vs Complaint (20.3%), eighth on the main Net Improvement board (9.71%), but 22nd on Steerability and 25th on Bash Recovery.
Elsewhere: GPT-5.6 Terra (Ultra) took TrackingAI's vision IQ test at 92.16, unseating Claude Opus 4 (Thinking)'s 87.5, and Claude Opus 4.8 opened the day's fourth debut board, the Vals AI Time Horizon Index — a Kerbal Space Program agent ladder — at 11.83% progress, between three and four missions into its 30-mission program, with GPT-5.5 trailing at 8.33.