OpenAI has released a large collection of mathematical results produced by an unreleased internal frontier model, making the material public through an open GitHub repository. The collection currently contains 722 manuscripts organized into 372 families of related results, spanning multiple mathematical disciplines. OpenAI says the work emerged from an evaluation program that moved toward open research problems after its existing mathematics benchmarks began to saturate.

The repository provides more than finished manuscripts. It includes source files, a catalogue of formalizations in Lean, ten abridged summaries of the model's reasoning and information about how the results were produced. OpenAI says the model was presented with roughly 4,000 problems and that the vast majority of accepted results were generated with a common procedure. On average, each published result used compute equivalent to about three hours of ChatGPT Pro thinking. Some results followed different procedures, and OpenAI identifies exceptions including work involving the Riemann zeta function and the Hodge conjecture for CM abelian varieties.

The scale of the release should not be confused with completed mathematical validation. OpenAI's own repository says the manuscripts are at different stages of verification, that not all have Lean formalizations and that some unformalized results could contain issues. Formal proof checking can establish that a proof satisfies a specified formal system, but the broader mathematical significance, novelty and correctness of hundreds of research claims still require scrutiny by specialists. OpenAI says it intends to preserve revision history and publish corrections as new versions rather than silently replacing earlier material.

An independent Advisory Group on Mathematics and Artificial Intelligence, whose members include prominent mathematicians and which operates independently of OpenAI, called the release an important event for mathematics but explicitly said its advisory role should not be interpreted as an endorsement of the results or of OpenAI's process. The group described publication as the beginning rather than the completion of human understanding and assessment. That distinction is important because the release is simultaneously a research artifact, an evaluation of an unreleased model and a large body of claims that the mathematical community now has to examine.

The broader significance is methodological. Frontier-model evaluation is moving beyond benchmark scores toward attempts at producing new scientific knowledge, where verification becomes a central bottleneck. OpenAI has exposed enough of the process to permit substantial external inspection, including manuscripts and many formal artifacts, but it has not released the model itself. The strongest conclusion today is therefore not that hundreds of open problems have definitively been solved. It is that an internal AI system generated an unusually large body of candidate mathematical research, some formally checked and some still awaiting validation, creating a substantial new workload and opportunity for the research community.

References