The Aggregate Digest logo

The Aggregate Digest

Archives
Log in
Subscribe
August 27, 2026

Zhipu's GLM-5.3 Flash arrived today at fourteenth of 689 on the best-available table, the strongest open-weights model there apart from Kimi K3

The Aggregate Digest — Thursday, August 27, 2026

Zhipu's GLM-5.3 Flash arrived today at fourteenth of 689 on the best-available table, the strongest open-weights model there apart from Kimi K3. Its case is tool use. Toolathlon scores whether an agent can pick, order and combine tools to finish realistic tasks, and GLM-5.3 Flash takes 78.4 percent on it, which is first of the 39 models LLM Stats has run through the benchmark and second of 55 on the self-reported board, behind Claude Opus 5's 80.6.

The next results are multimodal. MMVU asks video questions across several academic disciplines, and a self-reported 80.5 puts GLM-5.3 Flash second of 45 there, behind Gemini 3.7 Flash's 82.3. Artificial Analysis scores agents on 220 real work tasks spanning 44 occupations, with a shell and a browser to do them with, and GLM-5.3 Flash comes third of 217 entries on that evaluation.

It also outrates the API-only GLM-5.3, which has been on the table since 15 August, by 1741 to 1734. Kimi K3 keeps the open-weights crown at 1755. Claude Opus 5 is first overall at 1790.

Low-level and long-horizon work is where it falls away. On kernelbench.com's CUDA suite, which has agents optimise CUDA kernels and scores the percent of hardware roofline reached, GLM-5.3 Flash manages 4.45 percent and finishes last of eight, while Claude Opus 5 reaches 196.1. In LLM Stats' testing of Agents' Last Exam it is tenth of thirteen, 26.3 against GPT-5.6 Sol's 52.7. Fifty boards sit behind the rating so far, so it has room to move either way.

Read the full digest on The Aggregate →

Don't miss what's next. Subscribe to The Aggregate Digest:
← Newer Handing an agent a library of skills is worth up to nine points on Appwrite Arena, and the three models that led the board without skills are the onl… Older → Claude Opus 5 is now the best available model, after winning two agentic boards in this morning's Artificial Analysis refresh
aibenchmarks.dev
Twitter
Telegram
Powered by Buttondown, the easiest way to start and grow your newsletter.