The Aggregate Digest logo

The Aggregate Digest

Archives
Log in
Subscribe
July 20, 2026

AI Benchmark Digest — 2026-07-20

AI Benchmark Digest — 2026-07-20

View on AI Benchmark Hub

Daily

New Benchmarks (8)

  • SWE-Milestone (Milestone Score (%)): Claude Opus 4.8 (Max, 1M) leads with 51.84 across 23 models. SWE-Milestone evaluates coding agents on software-evolution itineraries mined from real repositories as milestone DAGs, requiring sustained multi-milestone implementation; scored by milestone completion with the best harness retained per model.
  • ErdosBench (Judge Average (0-4)): GPT-5.6 Sol (xHigh) leads with 3.12 across 7 models. Ulam\
  • MWS Vision Bench (Overall Score (%)): Cotype Pro 3 leads with 73.6 across 38 models. MTS AI\
  • VSI-Super-Wild (Overall Accuracy (%)): Gemini 3.1 Pro (Preview) leads with 44.36 across 14 models. VSI-Super-Wild (ECCV 2026) probes spatial supersensing in genuinely long in-the-wild videos — up to 4+ hours — with human-verified questions on motion, ordering, and continuous counting; scored by overall accuracy.
  • FutureX (Overall Score (latest week, %)): Qwen 3.7 Max leads with 49.33 across 15 models. FutureX is ByteDance\
  • WeaveBench (PassRate (%)): Claude Opus 4.7 leads with 35.1 across 10 models. WeaveBench tests computer-use agents on 114 long-horizon real-world tasks that require interleaving GUI clicks with shell and code in one trajectory, graded by a trajectory-aware agent judge; scored by pass rate.
  • SolarBench (Pass Rate (%)): Claude Fable 5 leads with 53.8 across 11 models. Maingen\
  • EdgeBench (Score @12h (134 tasks)): Claude Opus 4.8 leads with 50.89 across 5 models. ByteDance Seed\

New Scores From Top-10 Models (7)

  • GPT-5.5 on ForecastBench: 64.6 Overall Score (higher is better) (#65/236)
  • GPT-5.6 Sol on GBA Eval: 0.5259 Overall Emulator Score (#6/21)
  • GPT-5.6 Sol on MyPCBench: 55.4 Perfect Rate (%) (#2/8)
  • GPT-5.6 Sol on YapBench: 18.5 YapIndex (lower is better) (#1/108)
  • GPT-5.6 Sol on YapBench: 18.5 YapIndex (lower is better) (#1/108)
  • Kimi K3 on Gert Labs Rankings: 67.79 GScore (%) (#5/87)
  • Kimi K3 on YapBench: 300.7 YapIndex (lower is better) (#63/108)

New #1 Leaders (4)

  • Design Arena (Website) (Elo): Kimi K3 (1396.0) beat GPT-5.6 Sol (1352.0) by 44.0.
  • P2PCLAW Innovative Benchmark (Average Score): API-User (9.1) beat cajal-9B-v2-q8-v7-4-fixed (8.2) by 0.9.
  • Spider 2.0-Lite (Accuracy (%)): ktx (73.67) beat DivSkill-SQL (73.13) by 0.54.
  • ProphetArena (1 - Brier Score): ap-gemini-3.5-flash (0.9629) beat Kimi K2.5 (0.93) by 0.03.
Don't miss what's next. Subscribe to The Aggregate Digest:
← Newer AI Benchmark Digest — 2026-07-21 Older → AI Benchmark Digest — 2026-07-19
aibenchmarks.dev
Twitter
Telegram
Powered by Buttondown, the easiest way to start and grow your newsletter.