AI Systems Achieve Verifiable Mathematical Reasoning
Synthesis history
7 versions, newest first.
-
Version 7 2026-06-08 18:25 UTC · 84 items
The main addition is concrete year-over-year competition data: AI models that failed USAMO 2025 are now performing strongly on USAMO 2026 [^26381][^26382], providing empirical support for Rohan Paul's benchmark-saturati…
-
Version 6 2026-06-07 08:15 UTC · 79 items
The main addition is Rohan Paul's explicit prediction that competition math benchmarks are near obsolescence and that a model capable of a perfect IMO score will exist within one year [^26074] — an escalation from his p…
-
Version 5 2026-06-05 02:22 UTC · 74 items
Two substantive additions this pass. First, AlphaProof Nexus has moved from a name mention to a system with concrete enumerated results: it solved 9 open Erdős problems and 44 OEIS conjectures [21218][21768] — Zvi Mowsh…
-
Version 4 2026-05-26 08:25 UTC · 69 items
Two significant additions this pass. First, LessWrong and Hacker News evidence confirms both Google and OpenAI achieved gold-medal performance at IMO 2025 [21034][21035], complicating the prior framing that OpenAI was '…
-
Version 3 2026-05-25 10:23 UTC · 55 items
Three additions meaningfully develop the story this pass. First, Terence Tao's GitHub wiki tracking AI contributions to Erdős problems [19803] directly addresses the prior open question about his involvement—it confirms…
-
Version 2 2026-05-25 06:00 UTC · 45 items
Three significant additions expand the story this pass: (1) A Medium report claims three Erdős problems fell in seven days with Terence Tao personally verifying each—a claim that requires independent confirmation but wo…
-
Version 1 2026-05-23 02:45 UTC · 25 items
Three major AI research organizations—OpenAI, Harmonic, and Google DeepMind—have each produced AI systems capable of generating or verifying non-trivial mathematical proofs, with the most striking result being OpenAI's …