- OpenAI published 372 AI-generated math results on GitHub, saying most came from one prompt to one agent, at about three hours of compute each.
- Lean formal checks confirm logic, not whether a result is relevant or original, and 25 Fields Medal winners warn about the goals behind mass production.
- Leaders using AI output at volume should ask who checks it, with what tool, and what that tool cannot judge.
When an answer costs three hours of machine time, the answer stops being the scarce thing. The scarce thing becomes the person who can check it. Mathematics offers an early look at that problem, and any organisation that uses AI output at volume may face a version of it.
What OpenAI released
OpenAI has published 372 mathematical results from an internal frontier model, as reported by The Decoder. Each is meant to solve an open problem or make real progress on one. Topics include faster versions of widely used computer algorithms and progress connected to the Riemann hypothesis.
The results sit in a GitHub repository with revision logs and citations, not in academic journals. According to OpenAI, a single agent given a single prompt produced almost all of them. A few took repeat tries.
That is far cheaper than the company's earlier Navier-Stokes solution. OpenAI says that effort took 10,000 agents working together and millions of dollars of computing. It has been under formal review for weeks.
How the checking works
A good share of the proofs ship with versions written in Lean. Lean is a programming language built so that a computer can check each step of a mathematical argument. If the steps hold, the proof is logically sound. A person does not need to trust the author.
The reason, per The Decoder, is scale. Reviewing this many results by hand could be more than the math community can manage. OpenAI says more Lean versions are planned.
There is a limit. Lean can confirm that a result is correct. It cannot judge whether the result is relevant or original. A spell checker works the same way: it confirms the words are spelled right, not that the essay is worth reading.
Who carries the cost
The producer sets the pace. The reviewers absorb it. OpenAI consulted an advisory group at the Institute for Advanced Study, which includes Fields Medal winner Timothy Gowers. Beforehand, the company drew a line. The mathematicians could advise on how results are communicated. They could not advise on whether or how fast results are produced.
The company also did not publish its prompts. It shared only average compute costs, not per-problem figures. A reviewer therefore cannot easily see how hard each result was to obtain.
The community is divided. In an open letter, 25 Fields Medal winners said that solving problems is a stand-in for what mathematics really seeks: deeper understanding and insight. Their concern is that churning out true statements in bulk might damage the setting in which new ideas take root.
Gowers has cautioned that in one or two decades, mathematics could hold a vast body of papers with no human community that fully grasps them. Terence Tao has said young mathematicians should be trained with a focus on the human side, with tightly limited AI tool use.
The lesson for other fields
The lesson here is a general one, and it is analysis rather than a claim from the source. When generation gets cheap, review capacity sets the real speed limit. A formal checker helps, but it answers only the question it was built for.
The mathematicians' dispute points to a second problem. Even when a result is verified as correct, someone must still decide whether it matters. Any team that leans on automated checks faces the same gap.
Questions to put to your team
First, for each AI-generated output you rely on, who checks it, and how many items can they check per week? Second, is there an automated checker, and what exactly does it test? Third, what does it not test, and who covers that gap?
Fourth, who sets the production pace, and does that person bear any of the review cost? Fifth, can reviewers see how the output was made, or only the finished result?
OpenAI's repository will be judged by what mathematicians conclude about it. The broader point stands: producing answers is getting cheaper, and trusting them still takes people.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: The Decoder.





