OpenAI Publishes 722 AI-Generated Math Manuscripts Following Mathematician Backlash
OpenAI has released a repository of 722 formal mathematical proofs generated by an unreleased model to address skepticism from the academic community. The data dump details the compute costs and reasoning steps behind the system's previously announced solutions to complex problems.
By Tariq Nasser
When Google DeepMind published its AlphaGeometry system in 2024, the laboratory released both the model weights and the training code alongside its mathematical solutions. OpenAI has now taken a different route, publishing a massive repository of 722 formal mathematical manuscripts while keeping the artificial intelligence model that generated them strictly internal.[4]
The San Francisco-based company uploaded the collection of proofs and reasoning traces to GitHub on Tuesday, framing the release as a transparency measure. The repository contains solutions to problems drawn from a dataset of 4,000 mathematical challenges, complete with the step-by-step logic the unreleased system used to reach its conclusions.[1][7]
This data dump arrives after weeks of mounting pressure from the academic community. Mathematicians had criticized OpenAI for previously claiming its systems had solved unsolved mathematical problems without providing the underlying proofs to verify those assertions.[7]
"Sharing AI progress in mathematics requires more than just claiming a benchmark score," OpenAI stated in its official release notes. The company noted that the 722 manuscripts are intended to let domain experts evaluate the actual reasoning capabilities of its next-generation architecture before any commercial deployment.[1]
Verifying the reasoning traces
The newly published manuscripts are not simply final answers, but detailed formal proofs that can be checked by automated theorem provers like Lean. By formatting the output this way, OpenAI shifts the burden of verification from human peer review to deterministic software, bypassing the ambiguity that often plagues large language model outputs.[2][6]
According to the technical documentation accompanying the release, generating these 722 proofs required substantial computational resources. The company disclosed that the inference process cost approximately $2,000 in compute time, highlighting the massive energy and hardware requirements needed to produce formal mathematical logic at scale.[6]
The problems tackled by the unreleased model range from advanced combinatorics to algebraic geometry. While the company previously boasted about cracking famous open problems, this broader release demonstrates the system's ability to consistently format its reasoning across hundreds of distinct mathematical domains.[3]
However, the decision to withhold the model itself means independent researchers cannot test the system's failure modes or probe its limitations. Academics can read the 722 successful proofs, but they cannot see how many times the model hallucinated or failed before arriving at a verifiable answer.[4][5]
The academic reception
The initial reaction from the mathematical community has been a mix of validation and continued skepticism. While the formal proofs confirm that the model can indeed generate correct logic, the selective nature of the release leaves questions about the system's overall reliability unanswered.[7]
Researchers analyzing the GitHub repository have already begun feeding the manuscripts into automated theorem provers. Early reports indicate that the proofs are structurally sound, though some mathematicians note that the reasoning steps occasionally take inefficient or highly unconventional paths to reach standard conclusions.[2][5]
The release also serves as a strategic marketing maneuver for OpenAI's upcoming product cycle. By demonstrating advanced reasoning capabilities in a rigorous domain like mathematics, the company signals to enterprise clients that its next generation of models will handle complex, multi-step logic better than current iterations.[4]
"This is a preview of what inference-heavy models will look like," notes the analysis from MLQ.ai, pointing out that the system spent significantly more time computing each answer than standard conversational models do. The $2,000 compute cost for this batch underscores a shift toward spending processing power during the generation phase rather than just during training.[5][6]
Shifting the benchmark standard
For years, artificial intelligence laboratories have measured their progress using standardized multiple-choice tests or coding benchmarks. The publication of these 722 manuscripts suggests that formal mathematical proofs are becoming the new standard for proving that a model can actually reason rather than just retrieve memorized text.[3]
Because formal proofs can be mechanically verified, they eliminate the need for human graders and prevent models from succeeding through clever phrasing. If a proof compiles in a system like Lean, it is mathematically correct, regardless of whether a human or a machine wrote the underlying code.[2]
OpenAI's approach of releasing the outputs while hiding the engine contrasts sharply with open-source initiatives in the same space. While competitors publish their model weights to allow the community to build upon their work, OpenAI is treating its reasoning architecture as a proprietary trade secret.[4][7]
The company has not announced a timeline for when the model that generated these manuscripts will be available to the public or integrated into its commercial ChatGPT service. Until then, the 722 papers serve as the only concrete evidence of the system's capabilities.[1][5]
Evaluating the novelty
As mathematicians continue to parse the repository, the focus will likely shift from the correctness of the proofs to the novelty of the methods used. The true test of this unreleased model will be whether it has discovered genuinely new mathematical techniques, or simply automated the application of existing ones at an unprecedented scale.[3][6]
The dataset of 4,000 problems from which these 722 successful proofs were drawn indicates a success rate of roughly 18 percent on this specific, highly advanced benchmark. This metric provides a rare glimpse into the actual yield rate of current cutting-edge reasoning models when tasked with novel, graduate-level mathematics.[6][7]
By publishing the failures alongside the successes—or rather, by implicitly acknowledging the 3,278 problems the model could not formally prove—OpenAI provides a more grounded view of the technology's current state. The gap between the attempted problems and the published manuscripts defines the frontier of what artificial intelligence still cannot solve.[1][5]
The next phase of this development will depend on how the broader scientific community integrates these AI-generated proofs into human research. If mathematicians begin citing these 722 manuscripts as foundational steps for their own work, it will mark a significant transition in how mathematical knowledge is produced and verified.[3]
Key points
- OpenAI uploaded 722 formal mathematical manuscripts to GitHub, generated by an unreleased artificial intelligence model.
- The data release follows criticism from mathematicians regarding the company's previous undocumented claims of solving open problems.
- The inference process to generate the batch of proofs required approximately $2,000 in compute time.
- The company is withholding the model itself, preventing independent researchers from testing the system's failure modes.
Open questions
- When OpenAI plans to release the underlying model to the public or integrate it into commercial products.
- How many attempts or failed reasoning paths the model generated before arriving at the 722 successful proofs.
- Whether the system's training data included direct precursors to the specific problems it successfully solved.
Timeline
July 2026
OpenAI announces that an unreleased model, internally dubbed Astra, successfully solved 10 unsolved mathematical problems.
August 2026
Academic mathematicians publicly criticize the company for claiming benchmark victories without providing the underlying proofs for verification.
October 6, 2026
OpenAI publishes a GitHub repository containing 722 formal manuscripts and reasoning traces generated by the unreleased system.
- Open Science Advocates
- Argue that true scientific progress requires releasing model weights and training data alongside the results.
- AI Industry Analysts
- Focus on the compute costs and the strategic signaling of the release ahead of future product launches.
- Formal Verification Proponents
- Emphasize that the mathematical correctness of the proofs is what matters, regardless of how they were generated.
Perspectives this story doesn't cover
- Independent automated theorem proving developers
- Educational institutions adapting to AI-generated proofs
Sources
[1]OpenAIFormal Verification ProponentsSharing AI progress in mathematics
Read on OpenAI →
[2]Unite.AIFormal Verification ProponentsOpenAI Releases 722 Math Manuscripts From an Unreleased AI Model
Read on Unite.AI →
[3]QuartzFormal Verification ProponentsOpenAI's AI solved one famous math problem. Then it cracked hundreds more.
Read on Quartz →
[4]The Next WebOpen Science AdvocatesOpenAI publishes 722 maths papers written by a model it has not released
Read on The Next Web →
[5]MLQ.aiAI Industry AnalystsOpenAI Publishes 722 Math Manuscripts From an Unreleased AI Model
Read on MLQ.ai →
[6]Kingy.aiAI Industry AnalystsOpenAI's 722 Math Manuscripts: The Results, Proofs, Compute and Costs
Read on Kingy.ai →
[7]Times NowOpen Science AdvocatesAfter Mathematician Backlash, OpenAI Releases 722 Math Manuscripts With Proofs And Reasoning
Read on Times Now →
More in Technology
See all →Algorithmic Fairness
Why Equal Opportunity AI Metrics Hide False Positives That Equalized Odds Catches
7 sources
AI Data Routing
China Probes DeepSeek and Moonshot Over Routing Sensitive User Data to Anthropic's Claude
5 sources
Data Center Regulation
California Enacts Sweeping Resource Limits on AI Data Centers
5 sources
Vector Databases
The Curse of Dimensionality: Why Euclidean Distance Breaks Down in High-Dimensional Vector Databases
4 sources
Comments
Every angle. Every day.
Get Technology stories with full source coverage and perspective breakdowns, free every day.




