Rapid AI Benchmark Improvement: Small Models and New Entrants Closing Capability Gaps
Synthesis history
11 versions, newest first.
-
Version 11 2026-07-12 08:08 UTC · 225 items
Five new items this pass all relate to ARC-AGI-3 — the competition page, arxiv benchmark paper, and associated commentary — confirming the benchmark is now formally active rather than merely announced. None carried extr…
-
Version 10 2026-07-11 02:07 UTC · 220 items
The main addition is item 40065, Google's Android Bench update: GLM-5.2 is now included in Google's domain-specific coding benchmark alongside seven other models (Fable 5, Sonnet 5, Opus 4.8, Kimi K2.7 Code, MiniMax M3,…
-
Version 9 2026-07-08 02:27 UTC · 210 items
Jack Clark's Import AI 464[^39880] is the main substantive addition: it specifies Fable's KernelBench-Mega result (18.71X speedup, beating Claude Opus 4.8 at 14.4X and GPT 5.5 at 4.34X), reports the Remote Labor Index s…
-
Version 8 2026-07-04 18:42 UTC · 205 items
Lisan al Gaib's investigation into GLM-5.2's post-training methods[^39733] and Semgrep's independent cybersecurity benchmark result[^39731] are the two substantive additions this pass, deepening the distillation debate …
-
Version 7 2026-07-03 08:48 UTC · 182 items
EdgeBench from ByteDance Seed is the substantive new development: a benchmark measuring in-context experiential learning over 12–72 hour tasks finds top frontier models doubling their learning speed every three months, …
-
Version 6 2026-07-02 02:44 UTC · 166 items
Agents-A1 (35B) is a new entrant reported July 1 by Rohan Paul, claiming 1T-model-level performance by training on 45K-token verified task trajectories and distilling specialist teacher models — a different efficiency a…
-
Version 5 2026-06-30 18:41 UTC · 151 items
Simon Willison has published hands-on testing of Ornith-1.0's 35B MoE variant, confirming proficient agentic performance at 103 tokens per second and noting the clean Apache 2.0 licensing from the Gemma 4 base — adding …
-
Version 4 2026-06-29 02:32 UTC · 142 items
New items this pass are thin — mostly social media posts with no cited claims or key quotes. One developer (Justin, item 35818) confirms DeepSWE benchmark results match their hands-on experience with GLM-5.2, adding a f…
-
Version 3 2026-06-27 18:15 UTC · 136 items
GLM-5.2's 22.8% ARC-AGI-2 score received official verification from the ARC-AGI evaluation team[^35208], converting a self-reported result into an independently confirmed one. A widely-shared analysis combining the veri…
-
Version 2 2026-06-26 08:28 UTC · 123 items
Zvi Mowshowitz extended his GLM-5.2 analysis, now framing it as a potential 'DeepSeek moment' for open-source AI agents[^34125]—a stronger characterization than his prior commercially skeptical stance, though the two fr…
-
Version 1 2026-06-25 02:13 UTC · 106 items
Two Chinese AI labs released models in mid-June 2026 that have drawn attention for closing parts of the gap to closed frontier systems. GLM-5.2, from Zhipu AI, scores near Claude Opus 4.7 on traditional benchmarks and …