OpenAI posted 372 families of AI-generated mathematical results on Tuesday, Oct. 6, and withdrew three manuscripts Wednesday over a sign error. An unreleased internal model produced the results. As of Wednesday, about 42% of the catalogue's top-line results had machine-checked Lean proofs. OpenAI says some of the unformalized results could have issues, and the model remains unavailable for outsiders to rerun.
What Changed
- OpenAI posted 372 families of math results from an unreleased internal model on Oct. 6, then withdrew three manuscripts on Oct. 7 over a sign error.
- As of Oct. 7, 300 of 719 top-line results (about 42%) had machine-checked Lean proofs; OpenAI says some unformalized results could have issues.
- An OpenAI spokesperson said nearly every result came from a single prompt to a single agent, a claim no outside party can test because the model is unreleased.
- OpenAI published no prompts and only average compute, short of what the Institute for Advanced Study's advisory group asked for on Sept. 29.
AI-generated summary, reviewed by an editor. More on our AI guidelines.
The first corrections
The Oct. 7 revision history identifies the error in “Algebraicity of Weil classes on split abelian eightfolds.” The sign error invalidated a stabilization-trace cancellation argument and the construction used by two dependent manuscripts, “Algebraicity of Kuga–Satake Correspondences for K3 Surfaces” and “The rational Hodge conjecture for products of K3 surfaces,” which were also withdrawn.
OpenAI revised 14 other manuscripts Wednesday, repairing proofs and correcting statements, and updated references to revised companion manuscripts in another 13. Withdrawn versions remain accessible with notices explaining the gaps.
Tuesday's release contained 722 manuscripts; the catalogue now lists 719 manuscripts in 372 families. A family groups related manuscripts, including a principal result, companion arguments, consequences or alternative proofs. OpenAI's repository says the evaluation expanded after its existing mathematical tests saturated. The model was posed roughly 4,000 problems during the evaluation, using an average of about three hours of ChatGPT Pro thinking compute per result, OpenAI's figures show. By Wednesday, 300 of those 719 top-line results had Lean formalizations, which allow software to check a proof's logic.
The single-agent claim
An OpenAI spokesperson said nearly every result came from a single prompt to a single AI agent, though some might have required multiple attempts. Work on a Riemann zeta zero-free region and the Hodge conjecture for CM abelian varieties departed from the standard procedure; the zeta writeup was edited by a human for readability. The same spokesperson said many of the newly released results are not yet understood by the company's own mathematicians.
FREE WEEKDAY MORNING BRIEFING
Track what AI labs claim in math, and what holds up.
The Implicator Morning Briefing filters the AI news cycle to the stories worth your attention and explains their consequences. From San Francisco, every weekday at 4:45 a.m. Pacific, 7:45 a.m. Eastern.
About five minutes. No hype. No spam.
The same model produced the Navier-Stokes result announced Sept. 8, using a swarm of 10,000 agents rather than a single agent, at a compute cost of millions of dollars.
The single-agent and compute claims are OpenAI's own figures, and no outside party can rerun the model.
“Until and unless they release the model and people can replicate their results, I think you should treat any claims about one-shotting problems with a single agent as unverified,” said Andrew Sutherland, a mathematician at MIT.
Disclosure and human review
The Advisory Group on Mathematics and Artificial Intelligence, hosted at the Institute for Advanced Study, issued release recommendations Sept. 29. It asked labs to disclose model names and prompts, along with summarized reasoning and the time and estimated compute cost for each result. OpenAI published no prompts and gave average compute rather than figures for each result. Tuesday's release included ten abridged reasoning summaries.
Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.
The group also called for repositories outside AI labs' control and an end to testing advanced problems on proprietary models. This release came from a proprietary, unreleased model.
Daniel Litt, a mathematician at the University of Toronto, supported making the work public. “If we want to know the answers to these math questions, I see no reason why we should ask the company to keep them secret from us,” he said.
OpenAI plans to fund workshops, conferences and programs to help mathematicians understand the results, and is working to release the model.
“This release is the beginning, not the completion, of the process of human understanding,” the advisory group said Tuesday.
Frequently Asked Questions
What did OpenAI release?
On Oct. 6, OpenAI posted a GitHub catalogue of math results from an unreleased internal model. It now lists 719 manuscripts grouped into 372 families, after the release went out with 722 manuscripts.
Why were three papers withdrawn?
A sign error in "Algebraicity of Weil classes on split abelian eightfolds" invalidated an argument and a construction that two dependent manuscripts relied on. All three were withdrawn on Oct. 7, and 14 other manuscripts were revised.
How many results have been machine-checked?
As of Oct. 7, 300 of 719 top-line results, about 42%, had Lean formalizations, which let software check a proof's logic. OpenAI says some of the unformalized results could have issues.
Can anyone verify the single-agent claim?
Not yet. The model is unreleased, so outsiders cannot rerun it. MIT mathematician Andrew Sutherland said such claims should be treated as unverified until the model is released and results are replicated.
Did OpenAI follow the advisory group's recommendations?
Partly. The group asked for prompts and per-result compute costs and an end to testing advanced problems on proprietary models. OpenAI published no prompts, gave only average compute, and used a proprietary model.
AI-generated summary, reviewed by an editor. More on our AI guidelines.



Free AI briefing · Weekdays
The AI stories that matter, sourced and explained.
Join the Morning Briefing. It goes out every weekday at 4:45 a.m. Pacific, with later sends for the East Coast, Berlin and Tokyo.
Free when you sign up: The Paperclip Compendium, our tested guide to running AI agents.
Free. Unsubscribe in one click.
IMPLICATOR