The Aggregate Digest logo

The Aggregate Digest

Archives
Log in
Subscribe
August 30, 2026

Handing an agent a library of skills is worth up to nine points on Appwrite Arena, and the three models that led the board without skills are the onl…

The Aggregate Digest — Sunday, August 30, 2026

Handing an agent a library of skills is worth up to nine points on Appwrite Arena, and the three models that led the board without skills are the only ones in its top ten that skills did not help. Appwrite runs the same 24 models twice over the same Appwrite development tasks, once with its skills available and once without. The only variable between the two boards is the skills, and they raised 20 of the 24 scores.

How much a model gains runs opposite to how good it was on its own. Qwen 3.6 Plus goes from 87.4 without skills to 96.5 with, which is 21st place on the board to 10th. MiniMax-M2.7 gains 8.0 points, DeepSeek V4 Flash 7.5, Grok 4.6 6.2. At the top nothing moves: Claude Opus 5 scores 97.4 both ways, Claude Opus 4.8 loses 0.3, and Claude Fable 5, first without skills at 97.7, drops to 95.9 and 14th place with them.

Skills also compress the board. Without them the top ten spans 4.1 points; with them it spans 1.3. Muse Spark 1.2 took the with-skills lead today at 97.8, tied there with Grok 4.6 and a tenth of a point above GPT-5.5. Standing on the main table stops predicting the order: Qwen 3.6 Plus ranks 83rd of 1480 and costs $1.55 per task in our cost measurements, and with skills it finishes ahead of Claude Fable 5, which ranks ninth and costs $5.16.

Skill Coverage, new this week, asks the question head on: task success as a function of which skills an agent has available. The paper's answer is that more skills often do not help, and GPT-5.5 leads its five-model field at 53.26 percent. One board covers Appwrite's own domain and the other is five models wide, so neither is broad evidence. Both are worth reading before you wire a skill library into an agent.

Elsewhere. TimesFM-3 swept FEV-Bench, taking the average win rate at 85.93 to Chronos-2's 79.71 and first place on all four of the individual metrics. TimesFM-2.5, the previous generation, is ninth at 66.26 of 31 entries. Claude Opus 5 has held best available model since 21 August. Three new benchmarks are harder than their subject matter sounds: SaliTrap hides an obvious unstated condition in an everyday situation and Claude Opus 4.7 catches most of them at 62.5 percent; ORCA-bench hands a model the telemetry from an incident and asks it to name the root cause, where Claude Sonnet 4.6 leads at 30.9; PredAct-Bench injects noise into the tool outputs an agent depends on, and Gemini 3 Flash holds up best at 63.4.

Read the full digest on The Aggregate →

Don't miss what's next. Subscribe to The Aggregate Digest:
← Newer Claude Fable 5.1 arrived today and went straight to the top of the best available model table, ending Claude Opus 5's twelve-day hold at 1809 to 1786 Older → Zhipu's GLM-5.3 Flash arrived today at fourteenth of 689 on the best-available table, the strongest open-weights model there apart from Kimi K3
aibenchmarks.dev
Twitter
Telegram
Powered by Buttondown, the easiest way to start and grow your newsletter.