Oct 7 edition/Reporting & analysis
ResearchModelsPolicy

ResearchPapers, evidence & method

OpenAI posts 722 math manuscripts from an unreleased model on GitHub, but their verification is uneven

OpenAI released 722 AI-generated math manuscripts from an unreleased internal model, claiming results on major open problems. Verification is uneven, Lean coverage is partial, outside reproduction is impossible, and the release sidesteps an IAS-hosted advisory group's guidance on where to publish.

Sam Altman Speaks At OpenAI Developers Conference In San Francisco
Image: The Verge — Original article ↗
THE CORE IDEAS4 TAKEAWAYS
01

OpenAI says an unnamed internal model produced 722 manuscripts, grouped into 372 result families, after attempting about 4,000 problems. It reports that each result took compute roughly equal to three hours of ChatGPT Pro reasoning. The claimed results include a quasi-Riemann hypothesis zero-free region, the Unique Games Conjecture and the Hodge Conjecture for CM abelian varieties. All of these are vendor-reported, and no independent verification has been completed. [6] [1] [11]

02

Formal verification is partial. Unofficial counts put Lean formalization at 162 papers' main results or at 235 of the 372 families. The two figures count different units, and the formal code can cover less than an abstract claims. Lean shows that the logic is valid. It does not show that the formal statement matches the open problem, or that the result is new or significant. OpenAI itself warns that unformalized results may contain errors. [5] [10] [6]

03

The release is causing friction over governance. An IAS-hosted advisory group advised labs to publish through repositories no AI lab controls and to disclose the model, prompts and compute. OpenAI used its own GitHub account instead. Separately, 25 Fields Medalists argued that solving problems is not the same as understanding, and an earlier Navier–Stokes claim raised questions about provenance and credit. [1] [9] [3] [8]

04

Practitioners cannot use the capability itself. The model has no public access, API or model card, so nobody outside OpenAI can rerun it. The usable artifacts are the preprints, the source files, the Lean code and short reasoning summaries for 10 families. [6]

WHY IT MATTERS

OpenAI reports research-level results at about three Pro-hours each, far below September's 10,000-agent run. Implication: if the results hold, AI output could outpace mathematicians' capacity to review it.

Read the full assessment

Formal verification and human checks of formal statements would then become essential.

Executive brief

OpenAI has published 722 manuscripts, grouped into 372 "result families," on GitHub. An unnamed, unreleased internal model produced them. According to the OpenAI math repository, the model worked on about 4,000 problems, and the average result took compute equal to roughly three hours of ChatGPT Pro thinking. Results claimed in the batch include a "quasi-Riemann hypothesis" zero-free region, the Unique Games Conjecture and the Hodge Conjecture for CM abelian varieties. The repository says outright that results are "at different stages of verification" and that unformalized ones "could have issues." The release also departs from the IAS-hosted advisory group's guidance on where results should be posted (The Verge).

What changed and event timeline

  1. Astra ships

    OpenAI released GPT-6 Astra alongside 10 results on problems that had been open for a decade, according to.

  2. New model starts training

    OpenAI says it began training an internal model "significantly more capable than GPT-6 Astra" ().

  3. Navier–Stokes claim

    OpenAI said the model proved finite-time singularity formation for Navier–Stokes, with Lean formalization, using nearly 10,000 parallel agents over 88 hours. NYU's Tristan Buckmaster raised concerns about the provenance of the work ().

  4. Advisory group forms and issues guidelines

    IAS now hosts a nine-member Advisory Group on Mathematics and AI (AGMAI). It published release guidelines drawing on more than 600 survey responses ().

  5. The 722-manuscript release

    The papers went up in the Apache-2.0 openai/math repository, with protocols for revisions and citations (;).

Capabilities and access

  • Model: an unnamed, unreleased internal frontier model. There is no public access, no API and no model card (GitHub).
  • Procedure: OpenAI says almost every result came from a single prompt given to a single agent, under a fixed procedure (ExplainX).
  • What is public: the preprint PDFs with source files, Lean files, abbreviated reasoning summaries for 10 families, and BibTeX entries (GitHub).
Read the full section
  • Model: an unnamed, unreleased internal frontier model. There is no public access, no API and no model card (GitHub).
  • Procedure: OpenAI says almost every result came from a single prompt given to a single agent, under a fixed procedure (ExplainX).
  • What is public: the preprint PDFs with source files, Lean files, abbreviated reasoning summaries for 10 families, and BibTeX entries (GitHub).

Technical analysis for researchers and developers

  • Verification is uneven. The two counts measure different units (papers vs. families), and neither is official.
  • What Lean does and doesn't show. A Lean proof certifies that the logic is valid.
  • Not every result followed the standard pipeline. The repository lists the Riemann zeta zero-free region and the Hodge result as exceptions (GitHub).
Read the full section
  • Verification is uneven. Not every manuscript comes with Lean. Unite.AI counts 162 papers whose main result is formalized in Lean. ExplainX counts Lean links on 235 of the 372 families, but notes that the Lean coverage can be narrower than what the abstract claims. The two counts measure different units (papers vs. families), and neither is official.
  • What Lean does and doesn't show. A Lean proof certifies that the logic is valid. It does not show that the formal statement matches the open problem as posed, and it says nothing about novelty or significance.
  • Not every result followed the standard pipeline. The repository lists the Riemann zeta zero-free region and the Hodge result as exceptions (GitHub).
  • Reproducibility: nobody outside OpenAI can rerun the model. The roughly 4,000 problems attempted give a denominator for the results, but the selection criteria are not documented.

Claims and evidence

  • Batch size, compute and attempts (722 manuscripts, 372 families, about 3 hours per result, about 4,000 problems): vendor-reported (GitHub).
  • "Hundreds" of unresolved points solved: this description comes from AGMAI, not from OpenAI (The Verge). No independent audit of correctness has been reported.
  • Specific claimed results (quasi-RH, Unique Games, the free group factor problem): vendor-reported, with no completed independent verification (IBTimes Australia).
Read the full section
  • Batch size, compute and attempts (722 manuscripts, 372 families, about 3 hours per result, about 4,000 problems): vendor-reported (GitHub).
  • "Hundreds" of unresolved points solved: this description comes from AGMAI, not from OpenAI (The Verge). No independent audit of correctness has been reported.
  • Specific claimed results (quasi-RH, Unique Games, the free group factor problem): vendor-reported, with no completed independent verification (IBTimes Australia).
  • "90 of the top 500" open problems: commentary in Latent Space with no source document behind it.
  • Quality of earlier work: Timothy Gowers has said he would recommend at least one earlier proof to a top journal (Startup Fortune). That is a single expert view, not an assessment of this batch.

Context and prior work

This batch follows Astra's 10 results (August) and the Navier–Stokes claim (September). In that earlier work, about 10,000 agents ran for 88 hours on a single problem. The new batch reports roughly 3 hours per result (Express Tribune; Latent Space).

Read the full section

This batch follows Astra's 10 results (August) and the Navier–Stokes claim (September). In that earlier work, about 10,000 agents ran for 88 hours on a single problem. The new batch reports roughly 3 hours per result (Express Tribune; Latent Space). Anthropic and other rival labs are publishing math results too (The Verge). AGMAI asked labs to release through academic channels and to disclose the model, prompts and compute.

Limitations, safety and contested findings

  • OpenAI's own caveat: unformalized results "could have issues" (GitHub).
  • Disregarded guidance: AGMAI recommended posting in a repository that no AI lab controls. OpenAI used its own GitHub account instead (Verge).
  • Fields Medalists' letter: 25 Fields Medalists signed "A Severe Misalignment of AI in Mathematics."
Read the full section
  • OpenAI's own caveat: unformalized results "could have issues" (GitHub).
  • Disregarded guidance: AGMAI recommended posting in a repository that no AI lab controls. OpenAI used its own GitHub account instead (Verge).
  • Fields Medalists' letter: 25 Fields Medalists signed "A Severe Misalignment of AI in Mathematics." It argues that solving problems is only a proxy for understanding. Gowers warned that the literature could grow beyond what any human community can absorb (The Decoder).
  • Credit and provenance: OpenAI denies that it accessed Buckmaster's private work directly. It has said it cannot rule out benefiting from de-identified product data (Express Tribune).
  • Expected attrition: Will Depue expects some results not to survive scrutiny (Latent Space).

Business and practitioner implications

  • Pace and cost: if most results hold up, frontier models are producing research-grade math at about three Pro-hours per result.
  • Lean checking is a requirement now: for AI-generated proofs, Lean checking plus human review of the formal statements is becoming the minimum standard.
  • No access: practitioners can't use the capability itself. The only usable artifacts are the papers and the Lean code.
Read the full section
  • Pace and cost: if most results hold up, frontier models are producing research-grade math at about three Pro-hours per result. That is far cheaper than the agent swarms used weeks earlier.
  • Lean checking is a requirement now: for AI-generated proofs, Lean checking plus human review of the formal statements is becoming the minimum standard.
  • No access: practitioners can't use the capability itself. The only usable artifacts are the papers and the Lean code.
  • Reputational and legal exposure: labs face it over attribution and training-data provenance. Product terms that cover research done with user data deserve scrutiny.
  • Citation risk: citing these results before they are independently checked is risky. Track the repository's versioned corrections.

Sources

Read the full section
FOLLOW THE EVIDENCE

The source trail.

Sources (12)
A LITTLE LESS NOISE. A LOT MORE CONTEXT.

Stay curious.
Follow the evidence.

Independent perspectives, the original sources, and room for the questions that don't have easy answers.

How we build the brief