Supabase shipped its agent evals as a matched pair, with and without a skills pack — and the pack helps some models while hurting others
The Aggregate Digest — Monday, August 3, 2026
Supabase shipped its agent evals as a matched pair, with and without a skills pack — and the pack helps some models while hurting others. Both boards run the same 19 Supabase CLI build tasks (migrations, RLS, auth, data APIs), each model under its vendor-native harness, scoring the percent of tasks where every check passed. With the pack loaded, GPT-5.6 Sol takes a perfect 100 at both medium and low effort, and Claude Opus 5, Claude Sonnet 5 and Kimi K3 tie at 94.7. Strip it away and the order scrambles: Claude Sonnet 5 drops to 78.9, GPT-5.6 Sol at medium slips to 94.4, and Kimi K3 climbs to a perfect 100. A pack that costs one model 15.8 points and gains another 5.3 is not a uniform upgrade. Six models is a thin base to generalise from, but it is a cleaner controlled comparison than most boards offer.
Two crowns changed hands — the first movement since July 31, with August 1 and 2 both closing empty. DeepSeek V4 Flash (0731) took LLM Stats' NL2Repo at 54.2, displacing GLM-5.2's 48.9; the winner sits #40 overall against GLM-5.2's #36, so it is an upset in aggregate terms. Claude Opus 4.8 lifted MyPCBench's Perfect Rate to 62.0 from its own predecessor Claude Opus 4.6 at 58.2, with Qwen-CUA's 58.7 splitting the two — an in-family succession for a model that lost OckBench to Opus 5 three days earlier.
That same DeepSeek V4 Flash supplies the day's counterpoint: it tops NL2Repo and finishes last on Agents' Last Exam, its 25.2 trailing GPT-5.6 Sol's leading 52.7 by 27.5 points. That board arrives with a warning label — five rows, every one a vendor self-report, none verified upstream — so read the GPT-5.6 tier sweep (Sol 52.7, Terra 50.4, Luna 50.3) as claims rather than measurements.
Elsewhere. Hermes Best Models joined with 19 models over 25 agentic tasks and promptly showed why pass-rate boards saturate: Gemini 2.5 Flash and Gemini 3.1 Flash Lite both scored a perfect 100, while Claude Sonnet 4.6, GLM 5.1 and DeepSeek V4 Pro sit at 96 — the whole top five inside a single task. Gemini 2.5 Flash ranks #193 on our aggregate.