AI Policy Wiki
Dashboard

OpenAI Preparedness Framework V.2

high confidence · updated 2026-07-26

OpenAI's updated framework for evaluating and mitigating catastrophic risks from frontier models — 3 tracked categories (bio/chem, cyber, AI self-improvement), 4-level capability scoring, and Safety Advisory Group governance.

The Preparedness Framework is OpenAI's internal policy for tracking and preparing for frontier capabilities that create new risks of severe harm. Version 2, published April 15, 2025, defines "severe harm" as death or grave injury to thousands of people, or hundreds of billions of dollars in economic damage. It is a voluntary internal lab policy.

FieldValue
Publisher[[companies/openai\OpenAI]]
Version2 (updated from v.1)
DateApril 15, 2025
TypeVoluntary internal lab policy

Structure

The framework has four operational elements: (1) decide where to focus, defining which capability categories to track and at what thresholds; (2) measure capabilities, running frontier evaluations before and during deployment; (3) safeguard against severe harms, deploying safeguards sufficient to bring risk below threshold; and (4) build trust, through internal governance and external transparency.

OpenAI's launch announcement states that capabilities are prioritized for tracking against five criteria: the risk must be plausible, measurable, severe, net new, and instantaneous or irremediable (Source: openai.com). Evaluations combine a suite of automated, scalable tests — built to keep pace with a release cadence in which models improve without major new training runs — with expert-led "deep dives" to confirm the automated evaluations measure the right things (Source: openai.com).

Tracked categories

Three categories are tracked through every covered deployment. Biological and chemical capabilities carry a dual-use risk: the same capabilities that unlock discoveries and cures can reduce barriers to creating biological or chemical weapons. Models reaching a High threshold here cannot be deployed without sufficient safeguards.

Cybersecurity capabilities also carry a dual-use risk, spanning defensive capabilities for protecting systems and offensive capabilities enabling scaled cyberattacks and vulnerability exploitation. GPT-5.3 Codex was the first OpenAI model treated as High in Cybersecurity (Source: GPT-5.3-Codex System Card).

AI self-improvement capabilities address the risk that models capable of contributing to AI R&D could undermine human control of AI systems. This category replaced the prior "Autonomy" and "Persuasion" categories from v.1, reflecting the increasing relevance of recursive self-improvement concerns. Persuasion risks are handled outside the Preparedness Framework, including through the Model Spec, restrictions on the use of OpenAI tools for political campaigning or lobbying, and investigations into misuse such as influence operations (Source: openai.com).

Research Categories are a new addition in v.2: areas of capability that could pose severe harms but do not yet meet Tracked Category criteria, and for which OpenAI invests in developing threat models now. The announced focus areas are Long-range Autonomy, Sandbagging (intentionally underperforming), Autonomous Replication and Adaptation, Undermining Safeguards, and Nuclear and Radiological (Source: openai.com).

Capability threshold levels

Version 2 streamlines capability assessment to two operational thresholds, each mapped to specific commitments (Source: openai.com):

ThresholdDefinitionDeployment implication
HighCould amplify existing pathways to severe harmCannot deploy until safeguards sufficiently minimize the associated risk
CriticalCould introduce unprecedented new pathways to severe harmSafeguards required even during development

Correction note (2026-07-14): an earlier version of this page presented a four-level Low/Medium/High/Critical scoring table as the v.2 structure. OpenAI's launch announcement states that v.2 "streamlined levels to two clear thresholds" — High and Critical; the four-level scorecard belonged to the earlier framework structure that the update replaced (Source: openai.com).

Under a precautionary standard, if a model cannot be ruled out as crossing a High threshold, OpenAI treats it as High. Codex was treated as High Cybersecurity without definitive evidence because OpenAI could not rule it out (Source: GPT-5.3-Codex System Card).

Reports and frontier-landscape adjustments

Capability assessments are documented in Capabilities Reports (formerly the "Preparedness Scorecard"), and safeguard design and verification in dedicated Safeguards Reports; the Safety Advisory Group reviews both, assesses residual risk, and makes recommendations to OpenAI Leadership on deployment (Source: openai.com). The framework also provides that if another frontier AI developer releases a high-risk system without comparable safeguards, OpenAI may adjust its requirements — after rigorously confirming the risk landscape has changed, publicly acknowledging the adjustment, assessing that it does not meaningfully increase overall risk of severe harm, and keeping safeguards at a level more protective (Source: openai.com).

Internal governance: Safety Advisory Group

An internal cross-functional group of OpenAI leaders called the Safety Advisory Group (SAG) oversees the Preparedness Framework. SAG makes recommendations on safeguard levels, which OpenAI Leadership can approve or reject. The Board's Safety and Security Committee provides oversight. External advisors and US government partners also inform the process.

Differences from v.1

OpenAI cited four drivers for the update. Stronger models need more planning: prior model limitations provided a natural safety buffer, and that buffer is shrinking as models approach science-capable, agentic status. More frequent deployments, enabled by reasoning advances, drive faster deployment cycles that require scalable rather than only deep evaluations. A dynamic lab landscape, in which other labs pursue frontier work, makes coordination and knowledge-sharing more important. And accumulated experience from a year of safety-framework work at OpenAI and peer labs yielded clearer threat models and capability-elicitation techniques.

Key claims

#ClaimConfidenceSource notes
1Models reaching Critical threshold cannot be deployed even during development without safeguardsHighCore framework text
2OpenAI's evaluations had not shown models capable enough to pose severe bio/cyber risk without safeguards as of April 2025MediumDated claim; Codex (June 2025) was treated as High Cyber
3The framework is informed by external academic researchers, domain experts, and US government partnersHighFramework text

Comparison to Anthropic RSP

The Preparedness Framework and the Anthropic RSP share a core structure: capability-threshold evaluations, tiered responses, and pre-deployment gates. They differ in three respects. The RSP uses an ASL taxonomy (ASL-1 through ASL-4+), while the PF uses Low/Medium/High/Critical per category. RSP v3.1 explicitly covers CBRN, autonomy, and societal risks, while PF v.2 covers Bio/Chem, Cyber, and AI Self-Improvement. The RSP requires external evaluators (METR) by policy, while PF v.2 relies on the internal SAG with external consultation. See AI Safety Frameworks Compared: RSP, Preparedness, NIST RMF, and Safety Cases for the full comparison.

Relationships