AI Benchmark Digest — 2026-07-20
AI Benchmark Digest — 2026-07-20
Daily
New Benchmarks (8)
- SWE-Milestone (Milestone Score (%)): Claude Opus 4.8 (Max, 1M) leads with 51.84 across 23 models. SWE-Milestone evaluates coding agents on software-evolution itineraries mined from real repositories as milestone DAGs, requiring sustained multi-milestone implementation; scored by milestone completion with the best harness retained per model.
- ErdosBench (Judge Average (0-4)): GPT-5.6 Sol (xHigh) leads with 3.12 across 7 models. Ulam\
- MWS Vision Bench (Overall Score (%)): Cotype Pro 3 leads with 73.6 across 38 models. MTS AI\
- VSI-Super-Wild (Overall Accuracy (%)): Gemini 3.1 Pro (Preview) leads with 44.36 across 14 models. VSI-Super-Wild (ECCV 2026) probes spatial supersensing in genuinely long in-the-wild videos — up to 4+ hours — with human-verified questions on motion, ordering, and continuous counting; scored by overall accuracy.
- FutureX (Overall Score (latest week, %)): Qwen 3.7 Max leads with 49.33 across 15 models. FutureX is ByteDance\
- WeaveBench (PassRate (%)): Claude Opus 4.7 leads with 35.1 across 10 models. WeaveBench tests computer-use agents on 114 long-horizon real-world tasks that require interleaving GUI clicks with shell and code in one trajectory, graded by a trajectory-aware agent judge; scored by pass rate.
- SolarBench (Pass Rate (%)): Claude Fable 5 leads with 53.8 across 11 models. Maingen\
- EdgeBench (Score @12h (134 tasks)): Claude Opus 4.8 leads with 50.89 across 5 models. ByteDance Seed\
New Scores From Top-10 Models (7)
- GPT-5.5 on ForecastBench: 64.6 Overall Score (higher is better) (#65/236)
- GPT-5.6 Sol on GBA Eval: 0.5259 Overall Emulator Score (#6/21)
- GPT-5.6 Sol on MyPCBench: 55.4 Perfect Rate (%) (#2/8)
- GPT-5.6 Sol on YapBench: 18.5 YapIndex (lower is better) (#1/108)
- GPT-5.6 Sol on YapBench: 18.5 YapIndex (lower is better) (#1/108)
- Kimi K3 on Gert Labs Rankings: 67.79 GScore (%) (#5/87)
- Kimi K3 on YapBench: 300.7 YapIndex (lower is better) (#63/108)
New #1 Leaders (4)
- Design Arena (Website) (Elo): Kimi K3 (1396.0) beat GPT-5.6 Sol (1352.0) by 44.0.
- P2PCLAW Innovative Benchmark (Average Score): API-User (9.1) beat cajal-9B-v2-q8-v7-4-fixed (8.2) by 0.9.
- Spider 2.0-Lite (Accuracy (%)): ktx (73.67) beat DivSkill-SQL (73.13) by 0.54.
- ProphetArena (1 - Brier Score): ap-gemini-3.5-flash (0.9629) beat Kimi K2.5 (0.93) by 0.03.
Don't miss what's next. Subscribe to The Aggregate Digest: