The Information Machine

Rapid AI Benchmark Improvement: Small Models and New Entrants Closing Capability Gaps

Synthesis history

11 versions, newest first.

  1. Version 11 2026-07-12 08:08 UTC · 225 items

    Five new items this pass all relate to ARC-AGI-3 — the competition page, arxiv benchmark paper, and associated commentary — confirming the benchmark is now formally active rather than merely announced. None carried extr…

  2. Version 10 2026-07-11 02:07 UTC · 220 items

    The main addition is item 40065, Google's Android Bench update: GLM-5.2 is now included in Google's domain-specific coding benchmark alongside seven other models (Fable 5, Sonnet 5, Opus 4.8, Kimi K2.7 Code, MiniMax M3,…

  3. Version 9 2026-07-08 02:27 UTC · 210 items

    Jack Clark's Import AI 464[^39880] is the main substantive addition: it specifies Fable's KernelBench-Mega result (18.71X speedup, beating Claude Opus 4.8 at 14.4X and GPT 5.5 at 4.34X), reports the Remote Labor Index s…

  4. Version 8 2026-07-04 18:42 UTC · 205 items

    Lisan al Gaib's investigation into GLM-5.2's post-training methods[^39733] and Semgrep's independent cybersecurity benchmark result[^39731] are the two substantive additions this pass, deepening the distillation debate …

  5. Version 7 2026-07-03 08:48 UTC · 182 items

    EdgeBench from ByteDance Seed is the substantive new development: a benchmark measuring in-context experiential learning over 12–72 hour tasks finds top frontier models doubling their learning speed every three months, …

  6. Version 6 2026-07-02 02:44 UTC · 166 items

    Agents-A1 (35B) is a new entrant reported July 1 by Rohan Paul, claiming 1T-model-level performance by training on 45K-token verified task trajectories and distilling specialist teacher models — a different efficiency a…

  7. Version 5 2026-06-30 18:41 UTC · 151 items

    Simon Willison has published hands-on testing of Ornith-1.0's 35B MoE variant, confirming proficient agentic performance at 103 tokens per second and noting the clean Apache 2.0 licensing from the Gemma 4 base — adding …

  8. Version 4 2026-06-29 02:32 UTC · 142 items

    New items this pass are thin — mostly social media posts with no cited claims or key quotes. One developer (Justin, item 35818) confirms DeepSWE benchmark results match their hands-on experience with GLM-5.2, adding a f…

  9. Version 3 2026-06-27 18:15 UTC · 136 items

    GLM-5.2's 22.8% ARC-AGI-2 score received official verification from the ARC-AGI evaluation team[^35208], converting a self-reported result into an independently confirmed one. A widely-shared analysis combining the veri…

  10. Version 2 2026-06-26 08:28 UTC · 123 items

    Zvi Mowshowitz extended his GLM-5.2 analysis, now framing it as a potential 'DeepSeek moment' for open-source AI agents[^34125]—a stronger characterization than his prior commercially skeptical stance, though the two fr…

  11. Version 1 2026-06-25 02:13 UTC · 106 items

    Two Chinese AI labs released models in mid-June 2026 that have drawn attention for closing parts of the gap to closed frontier systems. GLM-5.2, from Zhipu AI, scores near Claude Opus 4.7 on traditional benchmarks and …