AI research
AI 'solved' famous maths problems. Read the asterisks.
The results are real. So are the caveats nobody screenshots.
The answer
AI advanced real maths in May 2026 — but humans and Lean verified every step.
Two claims, one week. OpenAI: an internal reasoning model disproved Erdős's 1946 unit-distance conjecture. DeepMind: AlphaProof Nexus solved nine open Erdős problems plus 44 conjectures. The dunk-tweet version — 'Google beat OpenAI nine to one' — is fun and almost entirely beside the point, because the two results aren't even the same kind of thing. One is a single flashy disproof awaiting review; the other is a stack of formally certified proofs. Counting them against each other is a category error dressed up as a leaderboard.
Where the asterisks live
OpenAI's proof is human-checked and not yet through formal peer review. Real mathematicians — Timothy Gowers among them — read it, which is genuine evidence, but expert eyes have missed subtle gaps before; that's the whole reason peer review exists in the first place. DeepMind's nine, by contrast, are machine-verified in Lean. That is the genuinely impressive part — and it also means the win comes from a tight generate-then-verify loop, not a model soloing genius in one heroic leap. The model writes proof steps in Lean's formal language; the compiler throws out anything that doesn't hold and feeds the error back for another try. The cleverness is real; so is the scaffolding holding it up, and honest coverage names both.
Put the two side by side and the 'who won' question dissolves:
| OpenAI | DeepMind | |
|---|---|---|
| The claim | One 1946 conjecture, disproved | Nine Erdős problems + 44 conjectures |
| Checked how | Humans (incl. Gowers) | Lean, step by step, automatically |
| Status | Peer review pending | Each proof formally certified |
| The asterisk | 'Solved' ≠ peer-reviewed yet | Brilliant — within a checkable box |
The scoreboard crowd is comparing a flashy single disproof against a pile of certified ones. Different sports.
Why it's not the singularity either
Here's the part the cynics get wrong, though. The fact that DeepMind leaned on Lean isn't a gotcha — it's the innovation. A loop where 'AI proposes, formal system filters' works only because, in maths, 'correct' is mechanically checkable. That's rare. Most of the world — strategy, taste, law, the messy middle of science — has no Lean to keep the model honest. So the achievement is sharply bounded: brilliant inside a checkable box, untested outside it. AlphaProof Nexus also builds on DeepMind's earlier olympiad prover, which means this isn't a system that woke up able to do research maths; it's years of narrow, deliberate engineering pointed at exactly the kind of problem it can verify. Genuinely hard, genuinely narrow — both at once.
Trust the vendor, not the timeline
Hassabis moved quickly to temper expectations, saying the system is 'still not AGI' even as it points toward a more practical role for AI in verified mathematical research.
The tell is that the people closest to it are the most careful. Hassabis said straight out the system is 'still not AGI'. When the vendor undersells and the timeline oversells, trust the vendor. AI is now a real instrument for parts of maths — narrow parts, with a human or a checker holding the pen — and pretending otherwise just sets up the next disappointment cycle. The hype merchants and the dismissers are both wrong, in mirror-image ways; the accurate read sits unglamorously in the middle.
OpenAI described the result as the first time a prominent open problem, central to a subfield of mathematics, has been solved autonomously by AI.
So file it correctly. A real milestone, a narrow one, and a preview of how AI actually gets useful in serious domains: not by being trusted, but by being checked. The headline 'AI solves famous maths' is true. The implied 'AI is now a mathematician' is not. The machinery in between is the only part worth screenshotting.
Frequently asked questions
So is AI better at maths than humans now?
Who actually 'won', OpenAI or Google?
If it used a proof checker, did the AI really do it?
What's the catch with OpenAI's result specifically?
Sources
- An OpenAI model has disproved a central conjecture in discrete geometry — OpenAI, 20 May 2026
- Advancing Mathematics Research with AI-Driven Formal Proof Search (AlphaProof Nexus preprint, arXiv:2605.22763) — Google DeepMind / arXiv, 21 May 2026
- OpenAI's milestone math breakthrough played to AI's strengths — Understanding AI, 22 May 2026
- Google DeepMind's AlphaProof Nexus Solves Erdős Problems as AI Math Race Moves Beyond Benchmarks — WinBuzzer, 26 May 2026