Zhipu's GLM-5.3 Flash arrived today at fourteenth of 689 on the best-available table, the strongest open-weights model there apart from Kimi K3
The Aggregate Digest — Thursday, August 27, 2026
Zhipu's GLM-5.3 Flash arrived today at fourteenth of 689 on the best-available table, the strongest open-weights model there apart from Kimi K3. Its case is tool use. Toolathlon scores whether an agent can pick, order and combine tools to finish realistic tasks, and GLM-5.3 Flash takes 78.4 percent on it, which is first of the 39 models LLM Stats has run through the benchmark and second of 55 on the self-reported board, behind Claude Opus 5's 80.6.
The next results are multimodal. MMVU asks video questions across several academic disciplines, and a self-reported 80.5 puts GLM-5.3 Flash second of 45 there, behind Gemini 3.7 Flash's 82.3. Artificial Analysis scores agents on 220 real work tasks spanning 44 occupations, with a shell and a browser to do them with, and GLM-5.3 Flash comes third of 217 entries on that evaluation.
It also outrates the API-only GLM-5.3, which has been on the table since 15 August, by 1741 to 1734. Kimi K3 keeps the open-weights crown at 1755. Claude Opus 5 is first overall at 1790.
Low-level and long-horizon work is where it falls away. On kernelbench.com's CUDA suite, which has agents optimise CUDA kernels and scores the percent of hardware roofline reached, GLM-5.3 Flash manages 4.45 percent and finishes last of eight, while Claude Opus 5 reaches 196.1. In LLM Stats' testing of Agents' Last Exam it is tenth of thirteen, 26.3 against GPT-5.6 Sol's 52.7. Fifty boards sit behind the rating so far, so it has room to move either way.