Byte
OpenAI Just Published 722 Math Proofs. Now Someone Has to Read Them.
OpenAI Just Published 722 Math Proofs. Now Someone Has to Read Them.
Mahmud Hasan
October 8, 2026
What actually landed on GitHub
On Tuesday, October 6, OpenAI posted 722 mathematics manuscripts to a public GitHub repository (openai/math), released under the Apache 2.0 license — grouped into 372 "families" around one main result each, with companion papers attached. The company says they span 17 areas of mathematics and include claimed solutions to 377 problems that were previously open.
For context: OpenAI fed roughly 4,000 unsolved research problems to an internal frontier model that no one outside the company can use. The results it considered significant enough to publish consumed, on average, compute equivalent to about three hours of ChatGPT Pro's "thinking" mode each. Most of the manuscripts are dated September 23 and 24 — a month of accumulated output dropped on the math community in a single afternoon.
Altman singled out four results. The quasi-Riemann hypothesis: a proof that the zeros of the Riemann zeta function stay at least a fixed distance from the edge of the region they can occupy — a step toward, not a solution of, the million-dollar Riemann hypothesis. The Unique Games Conjecture, proposed by Subhash Khot in 2002, which pins down how close efficient algorithms can get to optimal answers for hard problems like delivery routes and exam timetables. A case of the Hodge conjecture (another Millennium Prize problem) for one special family of shapes. And the free group factor problem.
Beyond those: in at least 44 cases the model claims to have disproved an existing conjecture. It claims a faster theoretical method for matrix multiplication — beating a record set in August by a team that included Google DeepMind researchers, though it would not speed up any computer you own. It extends the Fermat's-Last-Theorem line of work to number systems containing the square root of minus one. It says five colors are not enough to color the plane so that no two points exactly one unit apart share a color — tightening a 1950s puzzle whose answer was known to lie between 5 and 7. None of it is peer-reviewed; OpenAI says some results "could have issues" and promises corrections as new versions.
A Lean check proves less than it sounds — and more than critics admit
The most misunderstood detail of the release: Lean. OpenAI has attached machine-checkable Lean formalizations to many of the papers — nearly two-thirds of the result families by one count, 162 papers with a formalized main result.
But three different thresholds keep getting collapsed into one. First: is the result mathematically important? Second: is there a machine-checked formal proof of the central claim? Third: has the human mathematical community understood, stress-tested, and accepted it? A Lean check clears the second bar. It says almost nothing about the other two. Experts still have to confirm the computer checked the correct statement — that the formalization matches what a mathematician means by "the Unique Games Conjecture," not a subtly easier version of it.
Here's the awkward part: the Lean formalizations in this release were themselves written by AI, and OpenAI lists their review status as "unchecked." The chain of trust runs: AI wrote the proof, AI wrote the check of the proof, and humans are supposed to audit both. That's not fraud — it's what happens when proof production becomes cheap and the verification machinery (journals, referees, seminar talks) was built for a world where producing a proof was the scarce step. Verification is the bottleneck now, and the bottleneck is human.
Why mathematicians are furious
The anger didn't start this week. In August, OpenAI published ten results from the same model, and some mathematicians said parts of the work drew on earlier research without credit; OpenAI said it would update the papers. On September 8, the company claimed the model had solved the Navier–Stokes problem — before the proof had been independently verified. (That proof is not in the new release, and a credit dispute with NYU mathematician Tristan Buckmaster is still unresolved.) On September 11, 25 Fields Medal winners signed a declaration warning that the rush by AI companies to announce results was harming the discipline. On September 29, nine mathematicians at the Institute for Advanced Study issued guidelines asking AI companies to publish in journals and disclose how results were produced.
OpenAI has clearly been listening — it formed and consulted the Advisory Group on Mathematics and Artificial Intelligence (AGMAI) at the Institute for Advanced Study, and says its advice informed this release. But the gap between the letter and the spirit of those guidelines is where the fight is now. TensorFeed graded the release against the AGMAI guidelines and found: a problem count (about 4,000) and Lean formalizations for 162 papers, but no model name, no prompts, no dollar cost, and a repository on OpenAI's own GitHub — published the same day the community-run repository Hexagon invited lab submissions. The advisory group asked for a repository no lab controls. It got a repo on OpenAI's GitHub.
The deepest worry isn't procedure, though. It's what happens to a discipline when result production accelerates past the capacity to understand results. The r/mathematics threads this week make uncomfortable reading: postdocs saying thesis chapters on problems solved overnight have been "completely destroyed and made useless"; one researcher with two related preprints in preparation writing, simply, "fuck." A PhD student who spent three years on a problem an unreleased model solved in an afternoon didn't lose a race — the race was redefined.
The quieter, maybe smarter approach
While OpenAI was demonstrating what a giant general-purpose model can do across hundreds of problems at once, the Israeli startup doubleAI — founded by Amnon Shashua — announced that its system had made the first improvement in 38 years on a central problem in extremal graph theory: raising the girth coefficient of certain regular graphs from 4/3 to 8/5, closer to the theoretical maximum of 2, via a new algebraic construction proved for infinitely many degrees. The proof was also formalized in Lean.
The contrast is the interesting part. OpenAI is going for breadth: thousands of problems, hundreds of results, a firehose. doubleAI is going for depth: an AI system that operates as a specialist, a mathematical expert rather than a question-answering machine. One of those approaches will probably turn out to matter more for actual mathematics. It's not obvious which.
What's obvious is that the evaluation infrastructure is being invented in real time. OpenAI released ten shortened summaries of the model's reasoning — a window into how it found things — plus compute estimates per result and formalizations where it could. Every one of those choices is a guess at what a trustworthy release looks like when no human has read all 722 papers, including the people who published them. The community's answer, so far: read them slowly, check them carefully, keep score in public.
What to actually do with this
If you're a developer or technically inclined: clone openai/math and read the overview catalogue — the 41-page list is the fastest way to see the shape of the thing. To understand the verification question mechanically, install Lean and compile one of the formalizations — watching a machine check a proof teaches what "checked" does and doesn't mean. If you're supervising anyone who cites "AI proved X" in a meeting, the sentence you need is: posted is not proved, and formalized is not accepted.
Watch the corrections log. OpenAI committed to recording corrections as new versions with earlier versions left accessible — a small, boring, admirable design decision; an honest error log is the one part of this release no one can argue with. Whether 377 open problems are actually closed will be decided the old-fashioned way: one mathematician, one proof, one "wait, is this right?" at a time. The firehose is new. The readers are the same.
References
- OpenAI — "Sharing AI progress in mathematics" (October 6, 2026)
- ThePrint — "OpenAI drops mother lode of hundreds of math problem solutions" (October 7, 2026)
- Business Standard — "OpenAI shares summaries of math problems solved by its unreleased model" (October 7, 2026)
- The Times of India — "OpenAI now says it has solved more than 350 such problems" (October 7, 2026)
- Calcalistech — "OpenAI is solving decades-old math problems. An Israeli startup is taking a different route" (October 8, 2026)
- CellCog — "OpenAI's 722 AI Math Papers: What's Proved, What's Checked" (October 6, 2026)
- TensorFeed — "OpenAI Posted 722 AI-Written Math Papers to Its Own GitHub" editorial (October 7, 2026)
- AnandTech Forums — discussion with r/mathematics community reactions (October 7, 2026)
Comments
More in Artificial Intelligence

Apple's New CEO Just Made Himself Its Design Chief. He's an Engineer. That's the Point.
John Ternus visits Apple's design studio several times a week — Tim Cook came about once a month. Seven years after Jony Ive left, the CEO has made himself the company's design chief, and the bet is that one decision-maker beats a committee.
Read more
