The Information Machine

AI Systems Achieve Verifiable Mathematical Reasoning · history

Version 6

2026-06-07 08:15 UTC · 79 items

What

Two AI systems have produced expert-validated results in open research mathematics. OpenAI's reasoning model disproved the Erdős unit distance conjecture (open since 1946) [4], confirmed as a genuine milestone by Fields Medalist Tim Gowers and University of Toronto professor Daniel Litt [7]. Google DeepMind's AlphaProof Nexus solved 9 open Erdős problems and 44 OEIS conjectures [8][9][10], a volume Zvi Mowshowitz calls a landmark result that received almost no mainstream press. Competition-level benchmarks, where both systems already reached gold-medal performance at IMO 2025 [1][2], are now described by some observers as approaching saturation — Rohan Paul expects a model capable of a perfect IMO score within one year [15].

Why it matters

Expert mathematician validation now attaches to specific AI-produced results, not just benchmark scores. The volume of AI-resolved conjectures is rising faster than the mathematical community can independently verify them, a concern Terence Tao has named directly [18].

Open questions

  • AlphaProof Nexus solved 9 Erdős problems and 44 OEIS conjectures [8][9] — what are the specific problems, how do they compare in difficulty to the OpenAI unit distance disproof, and how has independent verification proceeded?

  • If competition math benchmarks near saturation within a year as Rohan Paul predicts [15], what evaluation frameworks exist for research-level AI mathematics that can distinguish degrees of difficulty beyond IMO gold?

  • Tao's concern about 'mass-produced mathematics at scale' [18] — has the mathematical community or any institution developed an institutional response to review-capacity pressure as systems resolve dozens of conjectures simultaneously?

  • The UW Math AI Lab's dataset of erroneous Lean proofs [14] challenges the trust model behind DeepMind's and Harmonic's formal-verification architectures — how are current systems addressing false proof generation within Lean pipelines?

Narrative

In 2025 and early 2026, AI systems reached levels of mathematical performance that had previously seemed distant. Google DeepMind and OpenAI both achieved gold-medal-level performance at the 2025 International Mathematical Olympiad [1][2][3] — the first AI systems to clear that threshold. Those results validated competition mathematics, where problems are pre-selected and success criteria are fixed. The more contested question has been whether AI can produce genuine research mathematics: open problems, no curated problem set, no guaranteed solution exists.

On that question, expert mathematician endorsement has accumulated around OpenAI's Erdős unit distance disproof. The result — a counterexample to a discrete-geometry conjecture open since 1946, produced by a general-purpose reasoning model with no formal-system scaffolding [4] — appeared in an arXiv preprint [5] and was acknowledged by combinatorialist Gil Kalai [6]. Fields Medal laureate Tim Gowers stated 'there is no doubt that the solution to the unit-distance problem is a milestone in AI mathematics' [7]; University of Toronto professor Daniel Litt called it 'the first example of a result produced autonomously by an AI that I find exciting in itself, as opposed to as a leading indicator' [7]. Google DeepMind's AlphaProof Nexus produced results at greater volume: the system solved 9 open Erdős problems and 44 OEIS conjectures [8][9][10] — a scale that Zvi Mowshowitz considers a landmark milestone but one that received almost no mainstream media coverage [9].

The methodological gap between the two approaches remains substantive. DeepMind's architecture pairs large language models with the Lean theorem prover, formally verifying each reasoning step [11][12]. OpenAI's Erdős counterexample came from a general-purpose model with no such scaffolding — it survived external mathematician review but carries no machine-checkable audit trail. A Google research paper found that breaking proof-writing into planning and step-by-step checking raised general LLM performance on formal math benchmarks from under 10% to 70% [13], suggesting structured decomposition may be the operative mechanism even without full Lean verification. The University of Washington Math AI Lab's dataset of erroneous Lean proofs [14] adds pressure on the formal-verification side: if the verification layer itself produces false proofs, formal-system grounding may be less reliable than the architecture implies.

The discourse around benchmarks is shifting. Competition math, once the primary measure of AI mathematical capability, is increasingly treated as near-obsolete by bullish observers — Rohan Paul has said he expects a model capable of a perfect IMO score within one year and that he will be 'disappointed' if that threshold is not reached [15]. Terence Tao, who actively curates AI contributions on his Erdős problems GitHub wiki [16] and has personally verified several AI-assisted results [17], has raised a counterpoint: 'more and more mass-produced mathematics at scale' [18] is a quality concern whose weight grows as systems announce dozens of conjecture resolutions simultaneously. Gary Marcus has continued auditing AI math headlines against underlying results [19] — a position now placed directly against Fields Medalist Gowers's explicit endorsement of the OpenAI result.

Timeline

  • 1946-01-01: Paul Erdős poses the unit distance conjecture in discrete geometry [4]
  • 2024-07-01: DeepMind's AlphaProof and AlphaGeometry 2 earn silver medal at IMO 2024 [30]
  • 2025-07-21: NYT reports Google AI wins gold at IMO 2025; OpenAI also claims gold — first AI systems to clear this threshold [1][2][3]
  • 2025-08-03: Xena Project publishes mathematical community round-up of AI performance at IMO 2025 [21]
  • 2026-02-01: The Atlantic publishes longform profile on how Terence Tao uses AI in his research [23]
  • 2026-05-20: OpenAI announces unreleased reasoning model disproved the Erdős unit distance conjecture [4][25]
  • 2026-05-21: Gil Kalai acknowledges disproof as 'Amazing'; widespread media amplification including The Guardian [6][31][32]
  • 2026-05-22: arXiv preprint on the disproof appears; formal verification debate surfaces in coverage [5][11]
  • 2026-05-23: Reports of three Erdős problems falling in seven days with Tao verifying each; Gary Marcus publishes critical review of AI math headlines [17][19]
  • 2026-05-24: Tao's GitHub wiki tracking AI contributions confirmed; Tao quoted on 'mass-produced mathematics at scale'; AlphaProof Nexus first mentioned [16][18][22]
  • 2026-05-28: Zvi Mowshowitz reports AlphaProof Nexus solved 9 Erdős problems and 44 OEIS conjectures; notes landmark result with almost no mainstream media coverage [9][8]
  • 2026-06-01: Ars Technica: Fields Medalist Tim Gowers and Daniel Litt explicitly validate OpenAI's Erdős result as substantively exciting [7]
  • 2026-06-04: Google research paper: structured proof-writing raises general LLM formal math performance from under 10% to 70% [13]
  • 2026-06-07: Rohan Paul predicts competition math benchmarks near obsolescence and expects a model with a perfect IMO score within one year [15]

Perspectives

OpenAI

A general-purpose reasoning model with no mathematical specialization disproved an 80-year-old open conjecture and achieved gold-medal performance at IMO 2025, demonstrating discovery capability without formal-system scaffolding.

Evolution: Stronger: Fields Medalist Tim Gowers and Daniel Litt have explicitly validated the Erdős result as substantively exciting, moving expert endorsement from acknowledgment to enthusiasm.

Google DeepMind

Achieved gold at IMO 2025, maintains a Lean-grounded theorem-proving architecture where every step is formally verified, and AlphaProof Nexus has solved 9 open Erdős problems and 44 OEIS conjectures.

Evolution: AlphaProof Nexus has moved from a name mention to a system with concrete enumerated results; mainstream press coverage of those results remains thin despite their volume.

Terence Tao

Actively curates AI contributions on his Erdős problems GitHub wiki and has personally verified several AI-assisted results, while publicly expressing concern about 'more and more mass-produced mathematics at scale.'

Evolution: Consistent dual position as engaged participant and quality skeptic; the quality concern gains weight as systems resolve dozens of conjectures simultaneously.

Tim Gowers and Daniel Litt (mathematician validators)

Both explicitly confirm the OpenAI Erdős disproof as substantively significant — Gowers calls it 'a milestone in AI mathematics'; Litt calls it the first AI result he finds 'exciting in itself, as opposed to as a leading indicator.'

Evolution: Consistent since entering the thread; these remain the most senior mathematician endorsements of any specific AI math result to date.

Rohan Paul (AI commentator)

Competition math and competition coding benchmarks are approaching obsolescence; expects a model capable of a perfect IMO score within one year, and frames structured proof-decomposition as the key architectural mechanism behind AI's formal math progress.

Evolution: Escalated from reporting architectural improvements to making explicit timeline predictions about benchmark saturation — the most bullish non-industry voice in the thread.

Gary Marcus

Explicitly auditing whether AI math headlines match underlying results; skeptical of headline claims about OpenAI and Anthropic mathematical achievements.

Evolution: Consistent; now positioned against explicit Fields Medalist endorsement of the OpenAI result, making the skeptic-vs-validator contrast sharper.

Harmonic (Tudor Achim)

Formal verification is the key epistemological shift enabling trustworthy AI mathematics; AI could prove the Riemann Hypothesis by 2028.

Evolution: Consistent and promotional; no new substantive claims beyond the founding thesis.

Mathematical community (Zvi, Xena Project, LessWrong, Hacker News)

Fragmented: Zvi notes that landmark results like AlphaProof Nexus receive almost no mainstream coverage despite their magnitude; Xena Project provides formal-verification-community context; LessWrong and Hacker News probe what verification actually guarantees.

Evolution: Consistent; the coverage-gap observation remains a notable feature of the story.

Tensions

  • OpenAI's Erdős disproof came from a general-purpose model with no formal-system grounding; DeepMind's architecture requires every reasoning step to be Lean-verified — both have produced expert-validated results, but they represent different claims about what makes AI mathematical output trustworthy. [4][11][12][5][2][3][7]
  • Gary Marcus explicitly audits AI math headlines against underlying results [19], while Fields Medalist Tim Gowers and Daniel Litt have directly endorsed the OpenAI Erdős result as substantively exciting [7] — the most direct clash between the skeptical and endorsing positions. [19][7][4]
  • Tao observes 'more and more mass-produced mathematics at scale' [18] as a quality concern; Rohan Paul frames high-volume proof-search as precisely the mechanism for mathematical progress and expects benchmark saturation within a year [15] — the same phenomenon read as a problem versus a feature. [18][15][13]
  • AlphaProof Nexus solved 9 Erdős problems and 44 OEIS conjectures — results Zvi calls a landmark milestone — but received almost no mainstream coverage [9], while OpenAI's single Erdős disproof generated sustained press attention; the coverage gap does not track the volume of results. [9][8][4][20]
  • The UW Math AI Lab's dataset of erroneous Lean proofs challenges the 'every step formally verified' claim central to DeepMind's and Harmonic's architectures — if the verification layer itself produces false proofs, formal-system grounding may be less reliable than claimed. [14][11][25][12]
  • Tudor Achim predicts AI could prove the Riemann Hypothesis by 2028 [25]; Rohan Paul expects a perfect IMO score within one year [15] — both bullish timelines sit against DeepMind's cautious published research and Marcus's ongoing skepticism. [25][15][11][19]

Sources

  1. [1] Google A.I. System Wins Gold Medal in International Math Olympiad — reactive:ai-formal-math-breakthroughs
  2. [2] OpenAI claims gold-medal performance at IMO 2025 | Hacker News — reactive:ai-formal-math-breakthroughs
  3. [3] Google and OpenAI Get 2025 IMO Gold - LessWrong — reactive:ai-formal-math-breakthroughs
  4. [4] 😸 OpenAI solved an 80-year math problem by... disproving it — The Neuron (2026-05-22)
  5. [5] Remarks on the disproof of the unit distance conjecture - arXiv — reactive:openai-erdos-math-breakthrough
  6. [6] Amazing: Erdős' Unit Distance Problem was Disproved! It was ... — reactive:openai-erdos-math-breakthrough
  7. [7] An OpenAI model solved a famous math problem that stumped humans for 80 years — Ars Technica AI (2026-06-01)
  8. [8] Google DeepMind's AlphaProof Nexus Solves 9 Erdős Problems and 44 OEIS Conjectures | KuCoin — reactive:ai-formal-math-breakthroughs
  9. [9] AI #170: Lack of Executive Order — Zvi's AI Roundups (2026-05-28)
  10. [10] DeepMind AlphaProof Nexus solves 9 Erdős problems - AI Weekly — reactive:ai-formal-math-breakthroughs
  11. [11] Google DeepMind's new paper. — Rohan Paul Twitter (2026-05-22)
  12. [12] @tomflex @prz_chojecki Sure! DeepMind built AI agents that pair LLMs (for generating ideas) with the Lean theorem prover... — reactive:ai-formal-math-breakthroughs (2026-05-24)
  13. [13] Another great paper from Google. — Rohan Paul Twitter (2026-06-04)
  14. [14] UW Math AI Lab Releases Erroneous Lean Proofs Dataset - LinkedIn — reactive:ai-formal-math-breakthroughs
  15. [15] "Pretty soon, competition math, competition coding, is not going to be interesting anymore. — Rohan Paul Twitter (2026-06-07)
  16. [16] AI contributions to Erdős problems · teorth/erdosproblems Wiki — reactive:ai-formal-math-breakthroughs
  17. [17] Three Erdős Problems Fell in Seven Days, and Terence Tao Verified ... — reactive:ai-formal-math-breakthroughs
  18. [18] “I do see more and more mass-produced mathematics at scale." — Rohan Paul Twitter (2026-05-24)
  19. [19] Checking the math behind OpenAI and Anthropic's latest headlines — reactive:ai-formal-math-breakthroughs
  20. [20] OpenAI announces AI's biggest math breakthrough yet — reactive:openai-erdos-math-breakthrough
  21. [21] AI at IMO 2025: a round-up - Xena Project - WordPress.com — reactive:ai-formal-math-breakthroughs
  22. [22] @tobiamure @Polymarket This is today's big news from Google DeepMind: their new AI agent (AlphaProof Nexus) autonomously... — reactive:ai-formal-math-breakthroughs (2026-05-24)
  23. [23] The Edge of Mathematics - The Atlantic — reactive:ai-formal-math-breakthroughs
  24. [24] Now that it's 2026, how is Terence Tao's prediction holding up? : r/math — reactive:openai-erdos-math-breakthrough
  25. [25] 😺 🎙️ PODCAST: Can AI Solve Math's Biggest Mystery? — The Neuron (2026-05-20)
  26. [26] [PDF] Aristotle: IMO-level Automated Theorem Proving - arXiv — reactive:ai-formal-math-breakthroughs
  27. [27] Aristotle from Harmonic just proved Erdos Problem #124 in Lean all ... — reactive:ai-formal-math-breakthroughs
  28. [28] Harmonic — reactive:ai-formal-math-breakthroughs
  29. [29] Gemini with Deep Think achieves gold-medal standard at the IMO | Hacker News — reactive:ai-formal-math-breakthroughs
  30. [30] Google DeepMind's AI systems, AlphaProof and AlphaGeometry 2 ... — reactive:ai-formal-math-breakthroughs
  31. [31] OpenAI makes breakthrough on 80-year-old maths problem — reactive:openai-erdos-math-breakthrough
  32. [32] OpenAI's internal model disproves Unit Distance Conjecture of Erdos — reactive:openai-erdos-math-breakthrough