A compute threshold is a quantity of training computation, usually expressed in floating-point operations (FLOP), that a legal instrument uses to decide which AI models fall inside a regulatory category. Thresholds function as a proxy: rather than defining the regulated class by capability, which is difficult to specify in advance and contested to measure, the drafter substitutes an observable input quantity that is easier to verify and known to the developer before a model is released. The two figures that recur across instruments are 10²⁵ FLOP, used by the European Union, and 10²⁶ FLOP, used by most United States federal and state drafting.
Compute thresholds are distinct from compute governance, which concerns control over the physical inputs to AI capability — semiconductors, data centers, and electricity — and operates through export controls, procurement, and industrial policy rather than through model classification.
Thresholds in current and proposed instruments
| Instrument | Jurisdiction | Category | Threshold | Effect |
|---|---|---|---|---|
| EU AI Act (Regulation 2024/1689) Art. 51(2) | EU | GPAI model with systemic risk | >10²⁵ FLOP cumulative training compute | Rebuttable presumption of high impact capabilities |
| Executive Order 14110 — Safe, Secure, and Trustworthy AI § 4.2 | US federal (rescinded) | Dual-use foundation model | 10²⁶ FLOP for dense models; lower for biology sequence models | Reporting to Commerce under the Defense Production Act |
| California SB 1047 — Safe and Secure Innovation for Frontier AI Models Act (enrolled + veto) | California (vetoed) | Covered model | >10²⁶ FLOP and >$100M training cost | Would have imposed safety and liability duties |
| California SB 53 — Transparency in Frontier AI Act | California (enacted) | Large frontier developer | 10²⁶ FLOP plus $500M revenue | Frontier AI framework, transparency reports, incident disclosure |
| New York RAISE Act (S. 8828) | New York | Frontier model | >10²⁶ integer or floating-point operations | Safety-protocol and incident-reporting duties |
California SB 1047 — Safe and Secure Innovation for Frontier AI Models Act (enrolled + veto) also captured derivatives, treating a fine-tune exceeding 3×10²⁵ operations and $10M in cost as a covered model in its own right. Before January 1, 2027 its threshold was fixed in statute; after that date it was to be revised by the California Government Operations Agency.
The compound-trigger pattern
The clearest drafting trend across US instruments is the pairing of a compute figure with a second, non-compute criterion. California SB 1047 — Safe and Secure Innovation for Frontier AI Models Act (enrolled + veto) paired 10²⁶ FLOP with a $100M training-cost floor; California SB 53 — Transparency in Frontier AI Act pairs the same compute figure with a $500M revenue test. Brookings estimates that approximately 5–8 entities are covered at the SB 53 thresholds, and identifies a regulatory-cliff dynamic in which the thresholds create a structural incentive for developers to stay just below them (California SB 53 — Transparency in Frontier AI Act, via FPF and Brookings — California SB 53 Compliance Analyses (Oct–Dec 2025)). The compound structure narrows the covered set to well-resourced developers and, in the revenue variant, decouples coverage from training scale alone.
The European approach differs in a way that matters for how the figure operates. Article 51(1) of the EU AI Act supplies two independent classification triggers: that the model has high impact capabilities evaluated on the basis of appropriate technical tools and methodologies including indicators and benchmarks; or that the Commission decides, on its own initiative or following a qualified alert from the scientific panel, that the model has equivalent capabilities or impact judged against the criteria in Annex XIII. The compute figure enters at Article 51(2) as a presumption rather than a definition — a model "shall be presumed to have high impact capabilities" when cumulative training compute exceeds 10²⁵ FLOP (EU AI Act (Regulation 2024/1689)). A developer may in principle rebut the presumption, and the capability-and-impact route at Article 51(1)(b) can capture a model that does not meet the compute figure at all. The US instruments by contrast make the compute figure definitional: a model either crosses it or does not.
The proxy problem
The recurring objection to compute thresholds is that a fixed FLOP line captures different capability levels in different years. Algorithmic efficiency improvements and hardware gains mean that the compute needed to reach a given capability falls over time, so a threshold set to identify the frontier at enactment progressively captures models further behind it — or, conversely, allows a capable model trained efficiently to fall below the line. The point remains a live question for Executive Order 14110 — Safe, Secure, and Trustworthy AI: the premise that 10²⁶ FLOP identifies the frontier was well supported at the time of the order and has become increasingly contested as compute-efficiency gains make smaller models capable.
The EU AI Act's own text anticipates the threshold aging. Article 51(3) empowers the Commission to amend the thresholds and supplement the benchmarks and indicators by delegated act "in light of evolving technological developments, such as algorithmic improvements or increased hardware efficiency" (EU AI Act (Regulation 2024/1689)). Commentary recorded on that page describes the 10²⁵ figure as an industry-lobbied compromise, notes that at the Act's passage only GPT-4 class models exceeded it, and observes that models trained with 10²⁴ FLOP may still pose systemic risks.
A second line of objection concerns the direction of the incentive the threshold creates. Governor Newsom's veto of California SB 1047 — Safe and Secure Innovation for Frontier AI Models Act (enrolled + veto) rested partly on compute-threshold grounds, and California SB 53 — Transparency in Frontier AI Act subsequently moved toward revenue-based triggers. The Brookings regulatory-cliff observation on SB 53 is the same objection in a different register: a bright-line trigger rewards designing a training run to land just underneath it.
Relation to capability categories
Compute thresholds are the operational substitute for capability definitions that the regulatory categories themselves resist. Foundation models are defined in the technical literature by adaptability across downstream tasks; frontier models are defined relative to the moving state of the art; and the EU's general-purpose AI model category substitutes observable training characteristics — broad data, self-supervision at scale, significant generality — for adaptability. In each case the compute figure enters where the drafter needs a line a regulator can check. A model can be a frontier model without meeting a given compute threshold in a given year, and can meet the threshold without being at the frontier.
Executive Order 14110 — Safe, Secure, and Trustworthy AI narrowed its category further by risk profile rather than by capability rank, restricting the dual-use foundation model to models whose capabilities bear on CBRN, offensive cyber, or loss of control, and pairing that with the 10²⁶ FLOP reporting figure. The order was rescinded in January 2025, but the definition and the threshold circulated into subsequent drafting, including California SB 1047 — Safe and Secure Innovation for Frontier AI Models Act (enrolled + veto) and the New York RAISE Act (S. 8828).
Thresholds in safety frameworks and safety cases
Compute figures also appear outside legal instruments, in developer safety frameworks and in the safety-case literature that reads them. Safety Cases for Frontier AI treats capability thresholds as proxy-based objectives — objectives that do not measure risk directly but measure other outcomes indirectly related to it — and expects early safety cases to use them while arguing that objectives should over time move toward direct measurements of risk. Its illustrative scope statement specifies pre-training with 10²⁶ total training FLOP, and its illustrative assumption block sets an explicit capability buffer covering one year of scaffolding and prompting improvements plus additional training using over 10% of pre-training compute, with a new safety case required once those buffers are surpassed. The paper is also explicit that a safety case resting on a framework's capability thresholds must justify the choice of thresholds rather than merely verify adherence to them.
Open questions
- Whether the Commission will exercise the Article 51(3) delegated-act power to move the 10²⁵ FLOP figure, and on what evidence of algorithmic or hardware efficiency, has not been resolved.
- Whether compound triggers pairing compute with revenue or cost hold up as training costs and revenue diverge across developers is not yet observable from the enacted instruments.
- How the threshold instruments treat a model whose capability is reached through post-training enhancement rather than pre-training compute is addressed only partially, through derivative rules such as California SB 1047 — Safe and Secure Innovation for Frontier AI Models Act (enrolled + veto)'s 3×10²⁵ fine-tune provision.
Relationships
- related: Foundation Models — the technical category for which the compute figure is a regulatory proxy.
- related: Frontier Models — defined relative to the moving state of the art rather than a fixed compute line.
- related: General-Purpose AI (GPAI) — the EU legal category into which the 10²⁵ FLOP presumption is embedded.
- related: Compute Governance — control of the physical compute stack, as distinct from compute as a classification trigger.
- instance-of: Scaling Laws — the empirical relationship between training compute and capability on which the proxy depends.
- depends-on: EU AI Act (Regulation 2024/1689) — Art. 51(2) rebuttable presumption at 10²⁵ FLOP and the Art. 51(3) amendment power.
- depends-on: Executive Order 14110 — Safe, Secure, and Trustworthy AI — the 10²⁶ FLOP dual-use foundation model reporting threshold.
- related: California SB 1047 — Safe and Secure Innovation for Frontier AI Models Act (enrolled + veto) — compound compute-and-cost trigger and the derivative fine-tune rule.
- related: California SB 53 — Transparency in Frontier AI Act — the enacted compute-plus-revenue trigger.
- related: New York RAISE Act (S. 8828) — the 10²⁶ operations frontier-model definition.
- supports: Safety Cases for Frontier AI — supplies the proxy-based objectives the paper expects early safety cases to use.