The Preparedness Framework is OpenAI's internal policy for tracking and preparing for frontier capabilities that create new risks of severe harm. Version 2, published April 15, 2025, defines "severe harm" as death or grave injury to thousands of people, or hundreds of billions of dollars in economic damage. It is a voluntary internal lab policy.
| Field | Value | |
|---|---|---|
| Publisher | [[companies/openai\ | OpenAI]] |
| Version | 2 (updated from v.1) | |
| Date | April 15, 2025 | |
| Type | Voluntary internal lab policy |
Structure
The framework has four operational elements: (1) decide where to focus, defining which capability categories to track and at what thresholds; (2) measure capabilities, running frontier evaluations before and during deployment; (3) safeguard against severe harms, deploying safeguards sufficient to bring risk below threshold; and (4) build trust, through internal governance and external transparency.
OpenAI's launch announcement states that capabilities are prioritized for tracking against five criteria: the risk must be plausible, measurable, severe, net new, and instantaneous or irremediable (Source: openai.com). Evaluations combine a suite of automated, scalable tests — built to keep pace with a release cadence in which models improve without major new training runs — with expert-led "deep dives" to confirm the automated evaluations measure the right things (Source: openai.com).
Tracked categories
Three categories are tracked through every covered deployment. Biological and chemical capabilities carry a dual-use risk: the same capabilities that unlock discoveries and cures can reduce barriers to creating biological or chemical weapons. Models reaching a High threshold here cannot be deployed without sufficient safeguards.
Cybersecurity capabilities also carry a dual-use risk, spanning defensive capabilities for protecting systems and offensive capabilities enabling scaled cyberattacks and vulnerability exploitation. GPT-5.3 Codex was the first OpenAI model treated as High in Cybersecurity (Source: GPT-5.3-Codex System Card).
AI self-improvement capabilities address the risk that models capable of contributing to AI R&D could undermine human control of AI systems. This category replaced the prior "Autonomy" and "Persuasion" categories from v.1, reflecting the increasing relevance of recursive self-improvement concerns. Persuasion risks are handled outside the Preparedness Framework, including through the Model Spec, restrictions on the use of OpenAI tools for political campaigning or lobbying, and investigations into misuse such as influence operations (Source: openai.com).
Research Categories are a new addition in v.2: areas of capability that could pose severe harms but do not yet meet Tracked Category criteria, and for which OpenAI invests in developing threat models now. The announced focus areas are Long-range Autonomy, Sandbagging (intentionally underperforming), Autonomous Replication and Adaptation, Undermining Safeguards, and Nuclear and Radiological (Source: openai.com).
Capability threshold levels
Version 2 streamlines capability assessment to two operational thresholds, each mapped to specific commitments (Source: openai.com):
| Threshold | Definition | Deployment implication |
|---|---|---|
| High | Could amplify existing pathways to severe harm | Cannot deploy until safeguards sufficiently minimize the associated risk |
| Critical | Could introduce unprecedented new pathways to severe harm | Safeguards required even during development |
Correction note (2026-07-14): an earlier version of this page presented a four-level Low/Medium/High/Critical scoring table as the v.2 structure. OpenAI's launch announcement states that v.2 "streamlined levels to two clear thresholds" — High and Critical; the four-level scorecard belonged to the earlier framework structure that the update replaced (Source: openai.com).
Under a precautionary standard, if a model cannot be ruled out as crossing a High threshold, OpenAI treats it as High. Codex was treated as High Cybersecurity without definitive evidence because OpenAI could not rule it out (Source: GPT-5.3-Codex System Card).
Reports and frontier-landscape adjustments
Capability assessments are documented in Capabilities Reports (formerly the "Preparedness Scorecard"), and safeguard design and verification in dedicated Safeguards Reports; the Safety Advisory Group reviews both, assesses residual risk, and makes recommendations to OpenAI Leadership on deployment (Source: openai.com). The framework also provides that if another frontier AI developer releases a high-risk system without comparable safeguards, OpenAI may adjust its requirements — after rigorously confirming the risk landscape has changed, publicly acknowledging the adjustment, assessing that it does not meaningfully increase overall risk of severe harm, and keeping safeguards at a level more protective (Source: openai.com).
Internal governance: Safety Advisory Group
An internal cross-functional group of OpenAI leaders called the Safety Advisory Group (SAG) oversees the Preparedness Framework. SAG makes recommendations on safeguard levels, which OpenAI Leadership can approve or reject. The Board's Safety and Security Committee provides oversight. External advisors and US government partners also inform the process.
Differences from v.1
OpenAI cited four drivers for the update. Stronger models need more planning: prior model limitations provided a natural safety buffer, and that buffer is shrinking as models approach science-capable, agentic status. More frequent deployments, enabled by reasoning advances, drive faster deployment cycles that require scalable rather than only deep evaluations. A dynamic lab landscape, in which other labs pursue frontier work, makes coordination and knowledge-sharing more important. And accumulated experience from a year of safety-framework work at OpenAI and peer labs yielded clearer threat models and capability-elicitation techniques.
Key claims
| # | Claim | Confidence | Source notes |
|---|---|---|---|
| 1 | Models reaching Critical threshold cannot be deployed even during development without safeguards | High | Core framework text |
| 2 | OpenAI's evaluations had not shown models capable enough to pose severe bio/cyber risk without safeguards as of April 2025 | Medium | Dated claim; Codex (June 2025) was treated as High Cyber |
| 3 | The framework is informed by external academic researchers, domain experts, and US government partners | High | Framework text |
Comparison to Anthropic RSP
The Preparedness Framework and the Anthropic RSP share a core structure: capability-threshold evaluations, tiered responses, and pre-deployment gates. They differ in three respects. The RSP uses an ASL taxonomy (ASL-1 through ASL-4+), while the PF uses Low/Medium/High/Critical per category. RSP v3.1 explicitly covers CBRN, autonomy, and societal risks, while PF v.2 covers Bio/Chem, Cyber, and AI Self-Improvement. The RSP requires external evaluators (METR) by policy, while PF v.2 relies on the internal SAG with external consultation. See AI Safety Frameworks Compared: RSP, Preparedness, NIST RMF, and Safety Cases for the full comparison.
Relationships
- supports: AI Safety Cases and Frameworks — documents a specific implemented safety framework
- related: Anthropic's Responsible Scaling Policy (Version 3.1) — peer framework; different taxonomy, similar logic
- related: GPT-5.3-Codex System Card — shows PF v.2 in practice (High Cybersecurity designation, precautionary approach)
- related: GPT-5.4 Thinking System Card — shows PF v.2 applied to reasoning model
- related: Frontier AI Safety Commitments (Seoul, 2024) — PF operationalizes Seoul Commitments II and IV
- related: AI Safety Frameworks Compared: RSP, Preparedness, NIST RMF, and Safety Cases — cross-framework comparison analysis
- related: AI Scheming — framework addresses scheming via self-improvement threshold