AI Benchmark Digest — 2026-07-16
AI Benchmark Digest — 2026-07-16
Daily
New Benchmarks (3)
- Epoch AI - Blueprint Bench 2 (Score): claude-fable-5_unknown leads with 38.61 across 17 models.
- Epoch AI - Proofbench (Score): gpt-5.6-sol_max leads with 77.0 across 43 models.
- Epoch AI - Enigma Eval (Score): gpt-5.4-pro-2026-03-05_unknown leads with 23.82 across 42 models.
New Models (1)
- Inkling — ELO 2701, #10/1474, above Claude Opus 4.8, below GPT-5.6 Pro Sol
- SEAL - AudioMultiChallenge: 56.64 (#1/32)
- SEAL - AudioMultiChallenge - Text Output: 56.64 (#1/17)
New Scores From Top-10 Models (2)
- Claude Fable 5 on SvelteBench: 100.0 Average pass@1 (%) (#1/149)
- Claude Opus 4.8 on SvelteBench: 100.0 Average pass@1 (%) (#4/149)
New #1 Leaders (3)
- GDP.pdf (Strict Pass Rate (%)): Claude Fable 5 (Thinking, Max) (70.3) beat GPT-5.6 Sol (30.7) by 39.6.
- SEAL - AudioMultiChallenge (Score): Inkling (56.64) beat Gemini 3 Pro (Preview) (Thinking) (54.65) by 1.99.
- SEAL - AudioMultiChallenge - Text Output (Score): Inkling (56.64) beat Gemini 3 Pro (Preview) (Thinking) (54.65) by 1.99.
Don't miss what's next. Subscribe to The Aggregate Digest: