🔍 Read the full analysis: The Future Of OpenAI’s AI Mathematics Is Still An Open Question on ThorstenMeyerAI.com
Get school and study supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
OpenAI published 722 manuscripts generated by an unnamed, unreleased model, spanning 372 families of mathematical results. The company has not established that the claims are correct; outside mathematicians must check them, and even valid proofs may not produce useful new methods or understanding.
OpenAI has published 722 mathematical manuscripts produced by an unnamed model it has not released, presenting claims across 372 families of results that include some of mathematics’ best-known open problems. The company and the material do not establish that those claims are correct: OpenAI CEO Sam Altman said they have not yet been confirmed by outside mathematicians, leaving verification and the work’s value to the field unresolved.
OpenAI’s post and accompanying GitHub repository say the manuscripts cover areas including number theory, geometry, topology, operator algebras, theoretical computer science and mathematical physics. They were selected from roughly 4,000 problems posed to the model; OpenAI filtered that pool for what it considered an appropriate level of significance. The average result used about three hours of ChatGPT Pro thinking compute, according to the source material. The manuscripts are published under the Apache-2.0 license.
The catalogue includes claims involving the Unique Games Conjecture, Hilbert’s tenth problem over the rationals, the isomorphism of nonabelian free group factors, a proposed zero-free region for the Riemann zeta function to the right of Re(s) = 11/12, and the Hodge conjecture for CM abelian varieties. These are claims in manuscripts, not independently established solutions. OpenAI’s README cautions that some results without formal proofs could have issues. Lean formalizations are included for many, but not all, results; the repository does not make all 372 families equally easy to check.
Only 10 abridged reasoning summaries accompany the 372 families. The source material says the Riemann write-up was edited by people for readability and identifies it and the Hodge result as exceptions to the standard process. The collection’s scale is striking, but selection and presentation were controlled by OpenAI: no outside group is reported to have chosen the problems or verified the full set before publication.
722 proofs, one question: will any of OpenAI’s AI mathematics actually lead anywhere?
An unreleased, unnamed model produced claimed proofs of results that would each define a career. Sam Altman calls them “claims not yet confirmed by outside mathematicians.” The real question isn’t whether it’s impressive. It’s whether answers nobody understands become discoveries anyone can build on.
Same day: Alon, Bloom, Gowers, Litt, Sawin post a digested, human-verified version. The model for success.
Connes rigidity counterexample challenged within a day — constructed groups fail the required condition. Three rival machine “counterexamples” from different labs now circulate.
~10,000 agents, 88 hours, est. ~$22M at retail. Priority dispute; 25 Fields Medalists sign “A Severe Misalignment” — not saying it’s wrong, saying it’s not understood.
Altman now hedges at announcement — a shift from September. Verification has barely started.
Humans extract the technique, write it up, build on it. This is where downstream discovery comes from.
The question is answered; nobody learns anything reusable. Closes a door without opening a field.
The proof breaks, or proves a statement that doesn’t match the conjecture as mathematicians mean it.
The Unique Games Conjecture is the clearest case. Results like the optimality of Goemans–Williamson for Max-Cut are proved assuming UGC. A correct proof converts them all — no understanding required. A zero-free strip for zeta works the same way for prime-distribution results. Free group factors, Kadison, Mahler would redirect whole programmes — but how depends on the method, which means digestion.
Technology. A Navier–Stokes blow-up proof doesn’t change how anyone designs aircraft; engineering turbulence models never depended on the answer. Near-term consequences are mathematical, not industrial. “AI will cure cancer next” skips several steps.
“Verification abundance, adjudication scarcity” — making proof-checking cheap doesn’t reduce the burden of deciding what’s true and what matters. 722 manuscripts land on a review system built for a trickle, filtered by a selection nobody outside OpenAI made.
Humans re-deriving results, like Alon–Gowers et al. in May
Other people’s work building on these manuscripts
How many unformalized results survive expert checking
Do the Lean statements match the real conjectures?
Do any survive peer review?
Some of it, yes — where a literature is waiting (UGC), a correct proof pays off immediately; where a proof carries a new technique humans digest, it can open a field. Most of it, probably not on its own: at 722 manuscripts with 10 reasoning summaries, the Four Colour pattern is the likely default unless mathematicians are funded and given time. And some will be wrong — OpenAI says so itself. It’s an industry pattern, not one company’s: the forced-Euler result came from an Anthropic researcher, and rival machine-generated Connes “counterexamples” circulate from different labs. The proofs arrived this week. The discoveries, if they come, will arrive at the speed of human understanding.
Verification Is Only the First Test
For mathematicians, a proof can matter for more than whether it settles a conjecture. Its arguments may introduce techniques that others can understand, reuse and extend. The source material contrasts Andrew Wiles’ proof of Fermat’s Last Theorem, whose methods contributed to later work, with the computer-assisted proof of the Four Colour Theorem, which settled a question but is presented as yielding less reusable theory. That comparison is an interpretation of mathematical impact, not a judgment already made about OpenAI’s manuscripts.
The Unique Games Conjecture illustrates why verification could have consequences beyond one paper. Many results in theoretical computer science are proved conditionally on it, including claims about the limits of approximation algorithms. If the manuscript’s claim were verified in a form accepted by specialists, researchers would need to examine which conditional conclusions change. At present, that is a possible consequence, not an established result: the conjecture has not been shown solved by the publication alone.
The distinction matters to readers because a large output count does not by itself measure scientific progress. A valid proof that researchers cannot interpret may settle a statement without helping them solve related problems. Conversely, a mistaken proof or a proof of a subtly different claim can attract attention without changing the field. The collection’s practical impact depends on independent checking and whether people can extract ideas from it.
mathematics problem solving software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A Mixed Record This Year
The release follows three major OpenAI mathematics announcements this year, according to the supplied source material, and those episodes show why the new claims need careful review. In May, the company’s model produced a counterexample to the Erdős unit-distance conjecture. Five mathematicians—Noga Alon, Thomas Bloom, Tim Gowers, Daniel Litt and Will Sawin—then published a human-verified account. That episode offers one possible route from machine output to accepted mathematical work: researchers translate and check the argument.
OpenAI’s August “Ten Advances” announcement had a more contested outcome. A claimed counterexample to Connes’s rigidity conjecture was challenged within a day, with critics arguing that the constructed groups did not meet a required condition. The source also reports that multiple machine-generated counterexamples to the same conjecture have circulated. Those claims illustrate why apparent agreement among generated outputs does not replace checking whether an argument addresses the precise mathematical statement.
In September, OpenAI announced a Lean-formalized result about finite-time blow-up in the Navier–Stokes equations, produced using about 10,000 concurrent agents over 88 hours, according to the source. That announcement prompted a separate dispute over priority and the purpose of using famous open problems as benchmarks. The source says 25 Fields Medalists signed a declaration titled “A Severe Misalignment of AI in Mathematics.” Their objection, as described there, focused on whether benchmark-driven solutions foster human understanding—not a finding that the proof was incorrect.
Which Claims Will Survive Review
Independent verification remains the main unknown. The supplied material does not report that outside mathematicians have checked all 722 manuscripts or confirmed any of the collection’s most prominent claims. It also does not specify a shared review process, timetable, or public accounting of which results have been accepted, challenged or withdrawn.
It is also unclear how many manuscripts have complete, machine-checkable formalizations and what those formalizations cover. A formal proof can help check that an argument follows within a specified system, but readers still need to know whether its assumptions and statement match the conjecture researchers care about. OpenAI’s selection of 4,000 problems down to 372 families is another limit: the catalogue does not show how many problems produced no result or how the company judged significance.
Finally, correctness would not settle the question of impact. The source material provides no evidence yet about whether the results will yield reusable methods, shift related research or remain difficult for mathematicians to interpret. Those outcomes may differ across the collection rather than resolve into one verdict on the model.
Outside Mathematicians Must Check
The next meaningful step is independent examination of individual manuscripts, including whether each argument is valid and proves the stated result. For claims with Lean formalizations, researchers can inspect the formal proof; for unformalized work, specialists will need to assess the written reasoning and its assumptions. The supplied source gives no deadline for those checks and no schedule for OpenAI to publish further reviews.
Readers should watch for named mathematicians or research groups to release detailed, reproducible assessments—not just broad reactions to the size of the catalogue. A checked result could then be reformulated in a way the field can use, as happened with the earlier Erdős-related work. Until such assessments appear, the manuscripts are a substantial set of mathematical claims, not a confirmed list of discoveries.
Key Questions
Did OpenAI prove the Unique Games Conjecture?
OpenAI published a manuscript claiming a result involving the Unique Games Conjecture. The supplied source says the claim has not been confirmed by outside mathematicians, so it should not be described as a verified proof.
How many manuscripts did OpenAI release?
The collection contains 722 manuscripts, arranged into 372 families of related results, according to OpenAI’s post and repository.
Has the full collection been independently checked?
The supplied material does not report a full independent review. OpenAI’s README also warns that some unformalized results could have issues.
Why might a correct AI proof still have limited impact?
A proof can settle a mathematical question without providing methods that researchers can reuse. Its longer-term value depends on whether mathematicians can understand, explain and build on the argument.
What should readers look for next?
Look for detailed assessments from outside mathematicians that identify which specific claims have been checked, what assumptions they use and whether the result matches the stated conjecture.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
