The 10^26 FLOP Threshold: How the US Government Monitors Frontier AI
Under the Defense Production Act, the Department of Commerce requires mandatory safety reporting for any AI model trained using more than 10^26 computational operations.
By Sofia Matos
- National Security Advocates
- Argue that strict compute thresholds and Defense Production Act reporting are essential to prevent adversaries from acquiring dual-use foundation models.
- AI Safety Researchers
- Warn that static compute thresholds are vulnerable to algorithmic efficiency improvements, meaning future dangerous models might be trained below the 10^26 FLOP line.
- Commercial AI Developers
- Emphasize that compute thresholds provide a clear, objective metric for compliance, preferable to vague capability-based definitions that create legal uncertainty.
- Open-Source Advocates
- Argue that export controls and reporting requirements tied to compute thresholds disproportionately harm open-weight models and consolidate power among a few well-resourced labs.
Perspectives this story doesn't cover
- International Trade Partners
- Hardware Manufacturers
What we don’t know
- Because Defense Production Act disclosures are confidential, it is unknown what specific vulnerabilities AI developers have reported to the federal government.
- It remains unclear how quickly algorithmic efficiency will improve, potentially allowing models trained below the 10^26 FLOP threshold to achieve frontier-level capabilities.
- The Department of Commerce has not publicly disclosed its internal capacity to audit or verify the red-team testing data submitted by AI developers.
The US Department of Commerce holds the authority to compel the world's largest artificial intelligence developers to hand over their cybersecurity protocols and safety testing data. Using powers derived from the Defense Production Act—a 1950 statute originally designed to ensure military supply chains during the Korean War—the Bureau of Industry and Security (BIS) requires mandatory reporting from any company training a model that crosses a specific computational line. The agency is currently finalizing the technical rules that will govern these disclosures for the next generation of AI systems, determining exactly what data the government can extract from private laboratories.[1]
That computational line is 10^26 floating-point operations (FLOPs). Established by Executive Order 14110 in October 2023, the metric serves as the US government's primary tripwire for frontier AI regulation, shifting oversight from voluntary corporate commitments to mandatory federal compliance. If a "dual-use foundation model" requires more than 10^26 computational operations to train, the developer must notify the federal government on an ongoing basis, share the results of adversarial red-team safety tests, and detail the physical security measures protecting the model's underlying weights.[1]
The legal mechanism enforcing this threshold is exceptionally rare for the software industry. The executive order explicitly invokes the Defense Production Act to mandate these disclosures. As the Federal Register outlines, the order requires "companies developing or demonstrating an intent to develop potential dual-use foundation models to provide the Federal Government, on an ongoing basis, with information, reports, or records regarding... any ongoing or planned activities related to training." This transforms the Department of Commerce from a trade promoter into a national security auditor for the artificial intelligence sector.[1]
To understand the sheer scale of 10^26 FLOPs, it is necessary to quantify the physical hardware required to reach it. A floating-point operation is a basic mathematical calculation, such as adding or multiplying two decimal numbers. The threshold of 10^26 represents 100 septillion operations. Reaching this level of compute requires tens of thousands of advanced graphical processing units (GPUs) running continuously for months, consuming massive amounts of electricity and water in the process.[2]
For historical context, OpenAI's GPT-3, which catalyzed the generative AI boom upon its release, required an estimated 10^23 FLOPs of compute to train. The federal reporting threshold is 1,000 times larger than that benchmark. Even the world's most powerful conventional supercomputers cannot easily breach this limit. The Oak Ridge National Laboratory's Frontier supercomputer, the first to achieve exascale performance, operates at roughly 10^18 FLOPs per second. A system of that size would need to run continuously for years to collectively reach 10^26 operations.[2]
Because single supercomputers cannot reach this scale efficiently, AI developers rely on massive, distributed computing clusters. Recognizing this, the Department of Commerce's rules specifically target the datacenters that house these clusters. The regulations define covered computing clusters as facilities with machines transitively connected by networking exceeding 300 gigabits per second and possessing a theoretical maximum performance greater than 10^20 operations per second. If a company acquires or possesses a cluster of this size, it must report its existence and location to the federal government.[1]
The strategic intent behind the 10^26 figure was to set the regulatory ceiling just above the capabilities of the models that existed when the order was drafted in late 2023. By placing the threshold at 10^26 FLOPs, the Biden administration effectively exempted the current generation of widely deployed commercial models from the strictest reporting requirements, while ensuring that the next generation of frontier models would be captured. It drew a line between the AI of today and the AI of tomorrow.[2][4]
This approach contrasts sharply with international regulatory frameworks, highlighting a fundamental divergence in how governments measure AI risk. The European Union's AI Act, which entered into force in 2024, classifies models trained on more than 10^25 FLOPs as posing "systemic risk," subjecting them to stringent evaluation, risk mitigation, and cybersecurity obligations. The EU threshold is exactly one order of magnitude lower than the US standard.
This approach contrasts sharply with international regulatory frameworks, highlighting a fundamental divergence in how governments measure AI risk.
By setting the US threshold at 10^26 FLOPs, the Department of Commerce created a significant regulatory gap between the two jurisdictions. As of mid-2026, industry trackers estimate that roughly fifteen major models sit between the 10^25 and 10^26 FLOP range. This means a substantial cohort of highly capable AI systems face strict, mandatory regulation in Europe but fall entirely below the Defense Production Act's reporting threshold in the United States.[4]
However, the 10^26 FLOP metric is not solely used for domestic safety reporting. The Bureau of Industry and Security has also weaponized the threshold for international export controls. Under rules finalized in early 2025, the US government imposed a global licensing requirement on the export of any closed-weight AI model trained on more than 10^26 computational operations. This marked an unprecedented expansion of export controls from physical hardware to intangible software weights.[1]
This export control mechanism, designated under Export Control Classification Number (ECCN) 4E091, requires developers to obtain a license before transferring the weights of a covered model to any destination worldwide. The rule applies a strict "presumption of denial" to license applications, effectively blocking the export of America's most powerful AI systems to foreign adversaries and ensuring that models crossing the 10^26 threshold remain under tight domestic control.[1]
Despite its widespread adoption in federal policy, the reliance on a static FLOP count has drawn intense criticism from AI safety researchers and legal scholars. The primary vulnerability of a compute-based metric is that algorithmic efficiency is improving faster than hardware scaling. A model trained in 2026 using 10^25 FLOPs might achieve the exact same capabilities as a model that required 10^26 FLOPs in 2024.[3]
"Because it would currently cost more than $100,000,000 to train a model using >10^26 FLOP, the addition of the cost threshold did not change the scope of the definition in the short term," notes an analysis by Law-AI. "However, the cost of compute has historically fallen precipitously over time in accordance with Moore's law." This dynamic guarantees that the 10^26 threshold will eventually capture cheaper, more common models unless the government actively revises the standard.[3]
This vulnerability in the compute-based definition was a central debate during the drafting of California's SB 1047 in 2024. State lawmakers attempted to "future-proof" their bill by including an epistemic capability threshold, covering any future model that "could reasonably be expected to have similar or greater performance" as a model trained on 10^26 FLOPs in 2024. The goal was to regulate the output capabilities rather than just the input compute.[3]
Ultimately, California Governor Gavin Newsom vetoed SB 1047, specifically citing the rigid compute thresholds as a primary reason for rejecting the legislation. The veto effectively killed the most prominent attempt to implement the 10^26 FLOP threshold at the state level, leaving the federal standard as the undisputed, albeit imperfect, benchmark for frontier AI regulation in the United States.[3]
The Department of Commerce retains the authority to lower the threshold if algorithmic efficiency renders the 10^26 figure obsolete. The original executive order mandates that the Secretary of Commerce regularly update the technical conditions for models and computing clusters subject to reporting. Yet, defining capabilities remains a profound regulatory challenge. While compute is an objective, measurable input—a developer knows exactly how many GPUs ran for how many hours—capabilities are subjective outputs that often emerge unpredictably during the training process.[1][3]
The evidence supporting the actual efficacy of the 10^26 FLOP threshold remains thin. Because the reporting requirements are strictly confidential under the Defense Production Act, the public has no visibility into what specific vulnerabilities AI developers have reported to the federal government. Furthermore, the Department of Commerce has not publicly demonstrated that it possesses the technical capacity or the specialized workforce required to rigorously audit the raw red-team testing data it receives from the world's leading AI laboratories.[4]
The next critical juncture for the 10^26 FLOP standard depends entirely on the deployment of the next generation of massive computing clusters. With companies building datacenters housing hundreds of thousands of next-generation GPUs, the physical infrastructure to routinely shatter the 10^26 threshold is now coming online. When those models finish training and trigger the Defense Production Act provisions, the Department of Commerce will receive its first true test of whether mandatory reporting can actually mitigate the risks of frontier artificial intelligence.[2]
Key points
- The US government uses the Defense Production Act to mandate safety reporting for AI models trained on more than 10^26 floating-point operations (FLOPs).
- This threshold is exactly one order of magnitude higher than the European Union's 10^25 FLOP baseline for systemic risk.
- The 10^26 metric effectively exempts the current generation of widely deployed commercial models while capturing next-generation frontier systems.
- The Bureau of Industry and Security also uses the 10^26 FLOP threshold to enforce global export controls on closed-weight AI models.
- Critics warn that static compute thresholds may become obsolete as algorithmic efficiency allows developers to train highly capable models with less computing power.
- 10^26
- US Defense Production Act FLOP threshold
- 10^25
- EU AI Act systemic risk FLOP threshold
- $100 million
- Estimated compute cost to reach 10^26 FLOPs
- 300 Gbit/s
- Datacenter networking speed triggering cluster reporting
How we got here
Oct 2023
President Biden signs Executive Order 14110, establishing the 10^26 FLOP reporting threshold under the Defense Production Act.
2024
The European Union AI Act enters into force, setting a lower 10^25 FLOP threshold for models posing systemic risk.
Sep 2024
California Governor Gavin Newsom vetoes SB 1047, rejecting an attempt to implement the 10^26 FLOP threshold at the state level.
Jan 2025
The Bureau of Industry and Security imposes global export controls on closed-weight models trained on more than 10^26 FLOPs.
Aug 2026
Industry trackers estimate that three frontier models have crossed the 10^26 FLOP threshold, triggering federal reporting requirements.
Sources
[1]Federal RegisterExecutive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence
Read on Federal Register →
[2]Independent InstituteThe Math Behind the AI Executive Order
Read on Independent Institute →
[3]Law and AIAI Safety ResearchersSB 1047 and the Challenge of Compute Thresholds
Read on Law and AI →
[4]Factlen Editorial TeamSynthesis by Factlen editorial team
Read on Factlen Editorial Team →
Comments
More in Artificial Intelligence
See all →AI Safety
Anthropic CEO Dario Amodei Calls for Coordinated AI Slowdown to Prevent Autonomous Botnet Threat
5 sources
Copyright Law
The Mechanics of the Fair Use Defense in Generative AI Training
6 sources
AI Architecture
The Four Components of a Retrieval-Augmented Generation (RAG) System: Indexing, Retrieval, Generation, and Evaluation
6 sources
Inference Economics
Explainer: Why the AI Market is Recalibrating Around 'Inference Economics' and Real-World ROI
4 sources
Every angle. Every day.
Get Artificial Intelligence stories with full source coverage and perspective breakdowns delivered to your inbox.




