Connect with us

NEWS

OpenAI’s 722 Math Papers Leave the Checking to Humans

OpenAI posted 722 math papers from a hidden model. Only some have Lean proofs, so the scarce work is now human checking.

Published

on

OpenAI on October 6, 2026 posted 722 mathematics manuscripts from an unreleased internal model, grouped into 372 result families. The company says the average result used the equivalent of three hours of ChatGPT Pro thinking, after the model was posed about 4,000 problems.

Finding candidate proofs is no longer the scarce step. Checking them, and then understanding them, is the work this dump hands to a field that cannot read 722 papers overnight.

OpenAI Put 722 Manuscripts on GitHub

The files live in a public GitHub repository under an Apache 2.0 license, with papers, Lean artifacts, and citation blocks in each directory. OpenAI’s catalogue lists 722 manuscripts in 372 families, and a family may bundle a main result with companion arguments, consequences, or another proof of the same claim.

The company says it consulted the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, and that it drew on that group’s public advice for this drop. The model that wrote the papers is still internal. OpenAI says it is working toward a responsible public release of that model and will keep testing it on mathematics and other sciences in the meantime.

Alongside the PDFs, the repo ships 10 abridged summaries of the model’s reasoning, estimates of compute in ChatGPT Pro units, and counts of problems tried. OpenAI also says it will fund workshops, conferences, and special programs so people can study major AI-written results, with details later.

A public folder is a better place to argue than a stage announcement, because specialists can open a file and attack a step. A URL is still not a stamp that a theorem is new, correct, or important. That gap is what the rest of this release is about.

How Many of the Papers Have a Lean Proof?

OpenAI’s own README says the collection sits at different stages of checking, that many manuscripts have Lean formalizations and many do not, and that it will add more computer-checked proofs as it gets them. Indexes of the repository’s formalization catalogue list a Lean-checked main result for 162 of the 722 papers, which leaves 560 without that machine check on the main claim.

Lean is a proof language a computer can verify, so a passing formalization is strong evidence that the encoded theorem follows from the encoded axioms. It does not, by itself, tell a reader whether the informal paper matches the Lean file, whether the result is new, or whether anyone understands why the argument works.

The README is blunt about the rest of the pile. “Some of the unformalized results could have issues,” it says, and the company says it will try to fix those quickly. Independent indexes of the overview split the 372 families across 17 fields. Theoretical computer science, combinatorics, and algebraic geometry hold the largest shares.

THE 372 FAMILIES BY FIELD

Field Result families
Theoretical computer science 40
Combinatorics 37
Algebraic and complex geometry 36
Number theory 31
Probability and statistical mechanics 29
Differential geometry 29
Mathematical physics 25
Operator algebras 19
Algebra 18
Topology 18
Real and complex analysis 16
Partial differential equations 16
Convex and metric geometry 15
Group theory 14
Dynamical systems and ergodic theory 12
Functional analysis 11
Mathematical logic 6

Those 17 rows add to 372 families, matching the README total. They are a map of where the model was pointed, not a scoreboard of theorems the community has accepted.

AGMAI Wrote the Rules a Week Earlier

The Advisory Group on Mathematics and Artificial Intelligence listed nine unpaid members on its site, among them Timothy Gowers, Martin Hairer, Edward Witten, and Melanie Matchett Wood. The Institute for Advanced Study hosts the group. It has no power inside any lab, and it formed after OpenAI approached some of the members about an external board; they built an independent body instead and invited the others in.

On September 29, 2026, seven days before the GitHub drop, the group published guidelines for the responsible release of AI-generated mathematics, informed by more than 600 survey replies. It asked labs to stop testing advanced problems on closed models. If they still release a pile of output no human yet understands, it said they must fund the later work of understanding, and that work should be led by the community, not by the lab that produced the files.

WHAT THE SEPTEMBER 29 NOTE ASKED LABS TO DO

  • Cite the literature: Search for related ideas and credit the papers that introduced them, even if the model found the same move on its own.
  • Use a community repo: Deposit results in a scholarly archive the lab does not control, with stable citations and a public record of later edits.
  • Show the receipts: For each result, name the model, publish prompts, a summarized chain of thought, time taken, and estimated compute cost.
  • Formalize when possible: Ship Lean artifacts to community standards, and state the formalization status clearly when a delay is unavoidable.
  • Record the misses: If many results go out at once, also say how many comparable problems the models tried and failed, and how the problems were chosen.

OpenAI met several of those asks in part. It published Lean files for many papers, kept version history, put BibTeX in each directory, and reported an average compute figure plus a count of problems posed. It did not name the model, did not publish per-result prompts and costs, and used its own GitHub rather than an archive it does not control, while saying it is still looking at community-hosted options that meet the group’s guidelines.

The September note also told labs to “refrain from treating the release of mathematical results as marketing vehicles to promote their models.” On October 6 the same group put a careful distance between advice and blessing.

Making this work public is a first step. This release is the beginning, not the completion, of the process of human understanding and the incorporation of the work into mathematical knowledge.

Advisory Group on Mathematics and Artificial Intelligence, October 6, 2026 statement

AGMAI said its role should not be read as a judgment of the results’ impact, or as an endorsement of how OpenAI obtained them. Only the mathematical community, it wrote, can do that assessment.

Zero-Free Regions and a Hodge Special Case

Most of the catalogue, OpenAI says, came from one fixed procedure on the unreleased model, with each result using about three hours of ChatGPT Pro thinking compute. Two lines of work sat outside that procedure: a zero-free region for the Riemann zeta function, and a proof of the Hodge conjecture for CM abelian varieties. The writeup for a zero-free region with real part greater than 11/12 was also edited by a person for readability.

That is not a proof of the Riemann hypothesis, and it is not the Hodge conjecture in full. It is a stronger zero-free strip than many classical bounds, and a Hodge result in a special geometric setting. The README’s 10 reasoning summaries name other headline families, including the irrationality exponent of π, the symmetric and general Mahler conjectures, Kaplansky’s direct-finiteness conjecture in characteristic two, the isomorphism of free group factors, and the three-dimensional relativistic Vlasov-Maxwell system.

Indexes of the Lean catalogue already mark some of those main claims as formalized, including a zero-free statement for Dirichlet L-functions with real part greater than 7/8, a counterexample to Kaplansky’s zero-divisor conjecture, and the symmetric Mahler case. The free-group-factor isomorphism is among the families that still lack a listed Lean main result. A public comparator run on the posted 7/8 zeta formalization reported that Lean’s default kernel accepted the encoded claim at that commit. That is a machine check of one file, not a community verdict on the paper.

OpenAI’s September 21 note, when the advisory group was announced, said the same internal model had “resolved more than 100 long-standing open problems across most areas of mathematics.” The October 6 catalogue is larger than that teaser: 372 families, which AGMAI described as solutions to “hundreds” of open questions. Those three figures count different things, and none of them is a referee report.

Checking Proofs Is Now the Scarce Work

The evaluation that produced this pile posed about 4,000 problems. OpenAI kept the outputs it judged significant, then grouped them into families. The average accepted result, the company says, used the equivalent of three hours of ChatGPT Pro thinking with the internal model. That is a cheap unit of discovery by the standards of a human research group, and it is why 722 manuscripts can appear on one Tuesday.

It is also why the bottleneck moved. A model can emit a convincing argument faster than a specialist can decide whether the argument is new, whether it cites the right ancestors, and whether the informal text matches the Lean, if any Lean exists. Within hours of the drop, independent browsers were already indexing titles, abstracts, fields, and Lean status so that people could sort the pile instead of scrolling a raw tree of folders.

THREE LAYERS OF ONE RESULT

  • Claimed: The model produced a proposed theorem and a writeup, and OpenAI judged it significant enough to keep.
  • Lean-checked: A main result has a formal proof in the catalogue that a kernel can accept, which is still only as good as the encoding.
  • Independently accepted: Specialists have read the argument, compared it with the literature, and the field treats it as part of mathematics.

Most of the 722 papers are still in the first layer. A minority have reached the second. The third layer is the one AGMAI says has not started. Journals, seminar speakers, and graduate students who were already living inside some of these questions now have to decide whether to drop a year of work, race to understand a machine proof, or wait for the first serious error report.

The Navier-Stokes Fight Still Shapes This Dump

The same unreleased model is the one OpenAI pointed at the Navier-Stokes problem in early September, after it began training the system on August 28, 2026. That earlier run used a swarm of about 10,000 agents, took 88 hours to a proposed singularity, and then took another 17 hours of Lean work, with about 2.7 million messages and 130 billion output tokens along the way. The Clay Mathematics Institute, which in 2000 named seven Millennium Prize Problems with a $1 million prize on each, said on September 11 that its own review process is deliberately unhurried.

That episode is why an advisory group exists at all. Credit fights, closed models, and announcement timing had already soured a large part of the research community on how labs talk about theorems. AGMAI’s September 29 text warned that proprietary testing “risks creating a two-tier system where labs outrun the rest of the field.” The October 6 dump is the first large test of whether a GitHub protocol and a promise of workshops can close that gap.

HOW THIS RELEASE WAS SET UP

  1. August 28, 2026: OpenAI begins training the new internal model it later points at open math problems.
  2. September 8, 2026: The company publishes its Navier-Stokes package from the same class of internal system, after an 88-hour agent run and 17 hours of Lean work.
  3. September 11, 2026: The Clay Mathematics Institute says it shares the field’s interest and that prize review will not be rushed.
  4. September 21, 2026: AGMAI is announced, and OpenAI says the model has resolved more than 100 long-standing open problems.
  5. September 29, 2026: AGMAI publishes release rules drawn from more than 600 survey replies, including a call to stop closed-model testing.
  6. October 6, 2026: OpenAI posts 722 manuscripts in 372 families on GitHub, with partial Lean coverage and 10 reasoning summaries.

The Hodge special case in this new catalogue sits on another Millennium line, again without a claim that the full prize problem is done. Clay’s $1 million rules still run on human time. GitHub does not.

Workshops Are Promised After the Papers

OpenAI says future drops will improve citations, exposition, and how results are presented so they are easier to understand. It also says it will fund workshops and special programs around major AI-written theorems. AGMAI had asked for exactly that kind of support, and had asked that nonprofit institutions, not the labs, decide who gets the money.

No dates, hosts, or budgets for those workshops are in the October 6 note. The model remains unreleased. Prompts for the 372 families are not in the README. The average compute figure is public; the per-result bill is not. Community-hosted mirrors are “being explored.”

AGMAI’s October 6 statement added a limit the GitHub folder cannot satisfy on its own. The future of mathematical research, it wrote, cannot consist only of understanding results produced by AI labs. Mathematicians still have to pose their own questions, and they need access to strong tools to do it. Until that access exists, 722 public papers are a library with a locked author and a long reading list.

Harry is the editor and lead writer of KERALANEWS 24X7, which he owns and runs as an independent publication. After ten years in journalism as a reporter and then an editor, he treats a story as something that keeps its history rather than a page that is silently replaced. When a report is updated, the new material is added with the time it arrived, and earlier text that turned out to be wrong is corrected in the open under the site's public corrections policy rather than deleted. Readers in any time zone can see how a story developed. Publishing around the clock never shortens the checking: the primary filing, statement, transcript or dataset is located first, and every number is confirmed against it before it appears. The site covers news, business and technology, science and sports, and entertainment, lifestyle and travel, with auto and gaming reported to the same standard, all for an international readership. Reader mail goes to Harry rather than to a form, at support@keralanews247.com.

Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending