NEWS
OpenAI’s 722 Math Papers Leave the Checking to Humans
OpenAI posted 722 math papers from a hidden model. Only some have Lean proofs, so the scarce work is now human checking.
OpenAI on October 6, 2026 posted 722 mathematics manuscripts from an unreleased internal model, grouped into 372 result families. The company says the average result used the equivalent of three hours of ChatGPT Pro thinking, after the model was posed about 4,000 problems.
Finding candidate proofs is no longer the scarce step. Checking them, and then understanding them, is the work this dump hands to a field that cannot read 722 papers overnight.
OpenAI Put 722 Manuscripts on GitHub
The files live in a public GitHub repository under an Apache 2.0 license, with papers, Lean artifacts, and citation blocks in each directory. OpenAI’s catalogue lists 722 manuscripts in 372 families, and a family may bundle a main result with companion arguments, consequences, or another proof of the same claim.
We’re releasing a broad range of new mathematical results produced by an internal frontier model.
We’ve been consulting with the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, and we have drawn on their advice and…
— OpenAI (@OpenAI) October 6, 2026
The company says it consulted the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, and that it drew on that group’s public advice for this drop. The model that wrote the papers is still internal. OpenAI says it is working toward a responsible public release of that model and will keep testing it on mathematics and other sciences in the meantime.
Alongside the PDFs, the repo ships 10 abridged summaries of the model’s reasoning, estimates of compute in ChatGPT Pro units, and counts of problems tried. OpenAI also says it will fund workshops, conferences, and special programs so people can study major AI-written results, with details later.
A public folder is a better place to argue than a stage announcement, because specialists can open a file and attack a step. A URL is still not a stamp that a theorem is new, correct, or important. That gap is what the rest of this release is about.
How Many of the Papers Have a Lean Proof?
OpenAI’s own README says the collection sits at different stages of checking, that many manuscripts have Lean formalizations and many do not, and that it will add more computer-checked proofs as it gets them. Indexes of the repository’s formalization catalogue list a Lean-checked main result for 162 of the 722 papers, which leaves 560 without that machine check on the main claim.
Lean is a proof language a computer can verify, so a passing formalization is strong evidence that the encoded theorem follows from the encoded axioms. It does not, by itself, tell a reader whether the informal paper matches the Lean file, whether the result is new, or whether anyone understands why the argument works.
The README is blunt about the rest of the pile. “Some of the unformalized results could have issues,” it says, and the company says it will try to fix those quickly. Independent indexes of the overview split the 372 families across 17 fields. Theoretical computer science, combinatorics, and algebraic geometry hold the largest shares.
THE 372 FAMILIES BY FIELD
| Field | Result families |
|---|---|
| Theoretical computer science | 40 |
| Combinatorics | 37 |
| Algebraic and complex geometry | 36 |
| Number theory | 31 |
| Probability and statistical mechanics | 29 |
| Differential geometry | 29 |
| Mathematical physics | 25 |
| Operator algebras | 19 |
| Algebra | 18 |
| Topology | 18 |
| Real and complex analysis | 16 |
| Partial differential equations | 16 |
| Convex and metric geometry | 15 |
| Group theory | 14 |
| Dynamical systems and ergodic theory | 12 |
| Functional analysis | 11 |
| Mathematical logic | 6 |
Those 17 rows add to 372 families, matching the README total. They are a map of where the model was pointed, not a scoreboard of theorems the community has accepted.
AGMAI Wrote the Rules a Week Earlier
The Advisory Group on Mathematics and Artificial Intelligence listed nine unpaid members on its site, among them Timothy Gowers, Martin Hairer, Edward Witten, and Melanie Matchett Wood. The Institute for Advanced Study hosts the group. It has no power inside any lab, and it formed after OpenAI approached some of the members about an external board; they built an independent body instead and invited the others in.
On September 29, 2026, seven days before the GitHub drop, the group published guidelines for the responsible release of AI-generated mathematics, informed by more than 600 survey replies. It asked labs to stop testing advanced problems on closed models. If they still release a pile of output no human yet understands, it said they must fund the later work of understanding, and that work should be led by the community, not by the lab that produced the files.
WHAT THE SEPTEMBER 29 NOTE ASKED LABS TO DO
- Cite the literature: Search for related ideas and credit the papers that introduced them, even if the model found the same move on its own.
- Use a community repo: Deposit results in a scholarly archive the lab does not control, with stable citations and a public record of later edits.
- Show the receipts: For each result, name the model, publish prompts, a summarized chain of thought, time taken, and estimated compute cost.
- Formalize when possible: Ship Lean artifacts to community standards, and state the formalization status clearly when a delay is unavoidable.
- Record the misses: If many results go out at once, also say how many comparable problems the models tried and failed, and how the problems were chosen.
OpenAI met several of those asks in part. It published Lean files for many papers, kept version history, put BibTeX in each directory, and reported an average compute figure plus a count of problems posed. It did not name the model, did not publish per-result prompts and costs, and used its own GitHub rather than an archive it does not control, while saying it is still looking at community-hosted options that meet the group’s guidelines.
The September note also told labs to “refrain from treating the release of mathematical results as marketing vehicles to promote their models.” On October 6 the same group put a careful distance between advice and blessing.
Making this work public is a first step. This release is the beginning, not the completion, of the process of human understanding and the incorporation of the work into mathematical knowledge.
Advisory Group on Mathematics and Artificial Intelligence, October 6, 2026 statement
AGMAI said its role should not be read as a judgment of the results’ impact, or as an endorsement of how OpenAI obtained them. Only the mathematical community, it wrote, can do that assessment.
Zero-Free Regions and a Hodge Special Case
Most of the catalogue, OpenAI says, came from one fixed procedure on the unreleased model, with each result using about three hours of ChatGPT Pro thinking compute. Two lines of work sat outside that procedure: a zero-free region for the Riemann zeta function, and a proof of the Hodge conjecture for CM abelian varieties. The writeup for a zero-free region with real part greater than 11/12 was also edited by a person for readability.
That is not a proof of the Riemann hypothesis, and it is not the Hodge conjecture in full. It is a stronger zero-free strip than many classical bounds, and a Hodge result in a special geometric setting. The README’s 10 reasoning summaries name other headline families, including the irrationality exponent of π, the symmetric and general Mahler conjectures, Kaplansky’s direct-finiteness conjecture in characteristic two, the isomorphism of free group factors, and the three-dimensional relativistic Vlasov-Maxwell system.
Indexes of the Lean catalogue already mark some of those main claims as formalized, including a zero-free statement for Dirichlet L-functions with real part greater than 7/8, a counterexample to Kaplansky’s zero-divisor conjecture, and the symmetric Mahler case. The free-group-factor isomorphism is among the families that still lack a listed Lean main result. A public comparator run on the posted 7/8 zeta formalization reported that Lean’s default kernel accepted the encoded claim at that commit. That is a machine check of one file, not a community verdict on the paper.
OpenAI’s September 21 note, when the advisory group was announced, said the same internal model had “resolved more than 100 long-standing open problems across most areas of mathematics.” The October 6 catalogue is larger than that teaser: 372 families, which AGMAI described as solutions to “hundreds” of open questions. Those three figures count different things, and none of them is a referee report.
Checking Proofs Is Now the Scarce Work
The evaluation that produced this pile posed about 4,000 problems. OpenAI kept the outputs it judged significant, then grouped them into families. The average accepted result, the company says, used the equivalent of three hours of ChatGPT Pro thinking with the internal model. That is a cheap unit of discovery by the standards of a human research group, and it is why 722 manuscripts can appear on one Tuesday.
It is also why the bottleneck moved. A model can emit a convincing argument faster than a specialist can decide whether the argument is new, whether it cites the right ancestors, and whether the informal text matches the Lean, if any Lean exists. Within hours of the drop, independent browsers were already indexing titles, abstracts, fields, and Lean status so that people could sort the pile instead of scrolling a raw tree of folders.
THREE LAYERS OF ONE RESULT
- Claimed: The model produced a proposed theorem and a writeup, and OpenAI judged it significant enough to keep.
- Lean-checked: A main result has a formal proof in the catalogue that a kernel can accept, which is still only as good as the encoding.
- Independently accepted: Specialists have read the argument, compared it with the literature, and the field treats it as part of mathematics.
Most of the 722 papers are still in the first layer. A minority have reached the second. The third layer is the one AGMAI says has not started. Journals, seminar speakers, and graduate students who were already living inside some of these questions now have to decide whether to drop a year of work, race to understand a machine proof, or wait for the first serious error report.
The Navier-Stokes Fight Still Shapes This Dump
The same unreleased model is the one OpenAI pointed at the Navier-Stokes problem in early September, after it began training the system on August 28, 2026. That earlier run used a swarm of about 10,000 agents, took 88 hours to a proposed singularity, and then took another 17 hours of Lean work, with about 2.7 million messages and 130 billion output tokens along the way. The Clay Mathematics Institute, which in 2000 named seven Millennium Prize Problems with a $1 million prize on each, said on September 11 that its own review process is deliberately unhurried.
That episode is why an advisory group exists at all. Credit fights, closed models, and announcement timing had already soured a large part of the research community on how labs talk about theorems. AGMAI’s September 29 text warned that proprietary testing “risks creating a two-tier system where labs outrun the rest of the field.” The October 6 dump is the first large test of whether a GitHub protocol and a promise of workshops can close that gap.
HOW THIS RELEASE WAS SET UP
- August 28, 2026: OpenAI begins training the new internal model it later points at open math problems.
- September 8, 2026: The company publishes its Navier-Stokes package from the same class of internal system, after an 88-hour agent run and 17 hours of Lean work.
- September 11, 2026: The Clay Mathematics Institute says it shares the field’s interest and that prize review will not be rushed.
- September 21, 2026: AGMAI is announced, and OpenAI says the model has resolved more than 100 long-standing open problems.
- September 29, 2026: AGMAI publishes release rules drawn from more than 600 survey replies, including a call to stop closed-model testing.
- October 6, 2026: OpenAI posts 722 manuscripts in 372 families on GitHub, with partial Lean coverage and 10 reasoning summaries.
The Hodge special case in this new catalogue sits on another Millennium line, again without a claim that the full prize problem is done. Clay’s $1 million rules still run on human time. GitHub does not.
Workshops Are Promised After the Papers
OpenAI says future drops will improve citations, exposition, and how results are presented so they are easier to understand. It also says it will fund workshops and special programs around major AI-written theorems. AGMAI had asked for exactly that kind of support, and had asked that nonprofit institutions, not the labs, decide who gets the money.
No dates, hosts, or budgets for those workshops are in the October 6 note. The model remains unreleased. Prompts for the 372 families are not in the README. The average compute figure is public; the per-result bill is not. Community-hosted mirrors are “being explored.”
AGMAI’s October 6 statement added a limit the GitHub folder cannot satisfy on its own. The future of mathematical research, it wrote, cannot consist only of understanding results produced by AI labs. Mathematicians still have to pose their own questions, and they need access to strong tools to do it. Until that access exists, 722 public papers are a library with a locked author and a long reading list.
-
BUSINESS2 months agoConsumer Sentiment Falls to 51.7 as Future Outlook Darkens
-
NEWS2 months agoPluto’s Heart Glacier Still Pushes Liquid Nitrogen Upward
-
GAMING1 month agoTake-Two Subpoenas Microsoft While GTA 6 Preorders Soar
-
ENTERTAINMENT1 month agoNetflix Weighs Hosting Peacock and Fox One in Its App
-
NEWS2 months agoUMMC Will Rebuild 30-Year-Old Cancer Labs With $2.4 Million
-
BUSINESS1 month agoAbbott Pays $670 Million to Exit a Missouri Food Verdict
-
NEWS1 month agoWater-Shedding Coatings Charge the Drops That Pierce Them
-
GAMING1 month agoGTA 6’s 80-Hour Run Counts Goals That Change the Story
