This page compares four approaches to governing how AI systems are developed, evaluated, and deployed, along five dimensions: what risk categories they cover, how they determine thresholds, how they handle accountability and verification, which risks they omit, and how they cope with rapidly changing capabilities. The four operate at different tiers: lab-level voluntary policies (Anthropic RSP and the OpenAI Preparedness Framework), a government-endorsed voluntary organizational framework (NIST AI RMF), and a conceptual/academic method (Safety Cases for Frontier AI). A fifth instrument, Anthropic's Frontier Compliance Framework (FCF), sits between lab-voluntary practice and binding law.
The comparison below is drawn from RSP v3.1 (April 2, 2026) and Preparedness v.2. The Anthropic policy has since been revised four times — v3.2 (April 29, 2026), v3.3 (May 26, 2026), and v3.4, effective July 8, 2026 — and the July 2026 version changes the accountability and verification analysis in ways set out under "The 2026 RSP revision" below (Anthropic Responsible Scaling Policy v3.4 (July 2026)).
The four frameworks at a glance
| Anthropic RSP v3.1 | OpenAI Preparedness v.2 | NIST AI RMF 1.0 | Safety Cases | |
|---|---|---|---|---|
| Developed by | Anthropic | OpenAI | NIST (US government) | Centre for the Governance of AI / GovAI |
| Type | Voluntary lab policy | Voluntary lab policy | Federal voluntary framework | Conceptual framework |
| Primary audience | Anthropic itself | OpenAI itself | Organizations deploying AI | Frontier labs + regulators |
| Risk categories | ASL-1 through ASL-4+ (autonomy, CBRN, societal) | Bio, Cyber, Autonomy, Persuasion | Govern, Map, Measure, Manage | Structured argument: claim → evidence |
| Trigger mechanism | Capability thresholds (evaluations) | Capability thresholds (scores per category) | Continuous organizational risk management | Safety case argument before deployment |
| Enforcement | Self-enforced | Self-enforced | None (voluntary) | Self-enforced / regulator-auditable |
Risk categories covered
The Anthropic RSP (v3.1) classifies systems by AI Safety Level. ASL-1 denotes no meaningful risk beyond existing technology; ASL-2 denotes some risk addressed through standard responsible deployment; ASL-3 is triggered by substantial autonomy uplift or meaningful CBRN (chemical/biological/radiological/nuclear) uplift and requires enhanced security plus evaluations; and ASL-4 covers autonomy risk (AI that could cause catastrophic harm without human action) or unprecedented CBRN uplift, requiring a demonstration of "no meaningful" risk before deployment. Its categories center on autonomy, CBRN, and societal impacts (Anthropic's Responsible Scaling Policy (Version 3.1)).
The OpenAI Preparedness Framework (v.2) uses four risk categories, each scored on four levels (Low/Medium/High/Critical): biological threats (uplift for biological weapons creation), cybersecurity (finding or exploiting vulnerabilities, creating offensive cyberweapons), autonomy (achieving complex goals without human oversight), and persuasion (political manipulation or large-scale influence). At the pre-deployment threshold, any Critical capability blocks deployment, while any High capability permits deployment with safeguards (OpenAI Preparedness Framework V.2).
The NIST AI RMF 1.0 is risk-category focused rather than capability focused, spanning a broader set of attributes: accuracy, reliability, explainability, privacy, fairness, accountability, transparency, and safety. It organizes risk management into four functions — Govern, Map, Measure, and Manage — and does not distinguish between frontier and narrow AI risks (NIST AI Risk Management Framework (AI RMF 1.0)).
Safety Cases are a structured argumentation method rather than a risk categorization. A safety case is an argument that "the system is safe enough to deploy," backed by evidence, an approach borrowed from aviation (DO-178C) and nuclear (IEC 61508). Its principal claim types are hard limits (the system never does X), soft limits (the system rarely does X under conditions Y), and relative claims (the system is safer than alternative Z) (Safety Cases for Frontier AI).
How thresholds are determined
Under the Anthropic RSP, capability thresholds are determined through structured evaluations run by Anthropic and independent evaluators (METR, Frontier Red Team). The Mythos System Card illustrates this in practice, with dozens of evaluations, expert red-teaming, virology trials, and autonomy benchmarks. Thresholds are defined qualitatively ("meaningful uplift") and interpreted through evaluation results. Because "meaningful uplift" is not precisely defined, different evaluators might reach different conclusions, and the RSP relies on Anthropic's judgment about which results constitute threshold-crossing.
The OpenAI Preparedness Framework is structured similarly: category scores at four levels (Low/Medium/High/Critical) are determined through specific evaluations. The GPT-5.3 Codex System Card documents a precautionary approach in which the model was treated as High Cybersecurity even without "definitive evidence" of threshold crossing, because OpenAI could not rule it out. Critical thresholds here are also judgment-based, and the framework acknowledges that evaluations may not capture real-world capabilities.
The NIST AI RMF sets no specific thresholds and instead provides for continuous risk management, with organizations defining their own acceptable risk levels within the framework. Relative to frontier labs, most organizations lack the expertise to define meaningful thresholds for advanced AI capabilities.
Safety Cases require quantifiable claims with evidence, such as "the system causes harm in fewer than 1 in 10^9 operational hours" with supporting evidence. This is the most rigorous of the four in its evidentiary demands, and correspondingly the hardest to apply when capabilities are poorly characterized.
Accountability and verification
| Self-assessed | External verified | Public reporting | |
|---|---|---|---|
| Anthropic RSP | Primarily | METR + Frontier Red Team (limited) | Partial (RSP text + system cards) |
| OpenAI Preparedness | Primarily | Third-party evaluators (limited) | Partial (system cards) |
| NIST AI RMF | Entirely | None built in | None required |
| Safety Cases | Proposed as external | Regulator-auditable | Could be public |
All current lab frameworks are primarily self-assessed. External evaluators such as METR and the MAIA Inspectorate exist but have limited capacity relative to the number of models being evaluated. The AI ethics auditing field is immature and currently insufficient for evaluating frontier model risks.
Risk categories the frameworks omit
Across the four frameworks, several risk types are addressed only partially or not at all:
| Risk | RSP | Preparedness | NIST | Safety Cases |
|---|---|---|---|---|
| AI labor/economic disruption | ✗ | ✗ | Partial | ✗ |
| Environmental impact | ✗ | ✗ | Partial | ✗ |
| Democracy/information manipulation | ✗ | Persuasion (partial) | Partial | Partial |
| Concentration of AI power | ✗ | ✗ | ✗ | ✗ |
| Non-human harms | ✗ | ✗ | ✗ | ✗ |
| Content moderation risks | ✗ | ✗ | Partial | ✗ |
The Taxonomy of Systemic Risks identifies 13 categories; the lab frameworks address four to five of them. The omitted risks are primarily social, economic, and political rather than technical safety, reflecting the labs' focus on catastrophic physical harm over structural and institutional harm.
Temporal dynamics
The frameworks are calibrated to current capability levels but must be applied to rapidly improving systems. The Anthropic RSP was designed with ASL-3 and ASL-4 as aspirational or future thresholds; Mythos Preview appears to be approaching or at ASL-3, and although the RSP process produced a restricted release, the speed at which capabilities cross thresholds is increasing. OpenAI's Preparedness Framework v.2 updates reflect capability changes since v.1, but GPT-5.3 Codex crossing High Cybersecurity under a precautionary classification implies the framework may lag capability assessments. NIST AI RMF 1.0 was published in January 2023; the AI Action Plan (July 2025) directed NIST to revise it to remove DEI, misinformation, and climate references, illustrating that voluntary frameworks are politically contestable. Safety Cases, as a conceptual framework, are the most adaptable across capability levels because the argument structure remains valid regardless of capability level; the difficulty lies in generating sufficient evidence.
The 2026 RSP revision
Version 3.4 of the Anthropic RSP, effective July 8, 2026, is the ninth published version and alters two of the dimensions compared above (Anthropic Responsible Scaling Policy v3.4 (July 2026)).
Unilateral commitments separated from industry recommendations. The prior RSP "committed to implementing mitigations that would reduce our models' absolute risk levels to acceptable levels, without regard to whether other frontier AI developers would do the same." Anthropic now argues that the relevant unit is ecosystem risk: "If one AI developer paused development to implement safety measures while others moved forward with training and deploying AI systems without strong mitigations, that could result in a world that is less safe—the developers with the weakest protections would set the pace." Section 1 of v3.4 therefore presents three columns — capability thresholds calling for heightened mitigations, Anthropic's own planned mitigations, and the industry-wide recommendation at each threshold — with the recommendations described as ones "we cannot commit to following… unilaterally." Because Anthropic states it "cannot presently give highly specific advance detail on what evaluations will determine whether risk thresholds have been passed," the recommendations are "structured around requiring analysis and arguments making a strong case for safety, rather than AI Safety Levels" — a move toward the argumentation structure of the Safety Cases method rather than the capability-tier structure of the RSP's own ASL scheme. Anthropic names the cost of that flexibility ("one actor's view of what constitutes good risk assessment and mitigation may be very different from another's") and states a preference for third-party governance over self-regulation, with harmonization across countries "to avoid a race to the bottom."
Competitor-contingent commitments. Appendix A commits conditionally where the main text does not commit unilaterally. If Anthropic has a clear lead, it will require a strong argument that catastrophic risk is contained and "will delay AI development and deployment as needed to achieve this, until and unless we no longer believe we have a significant lead." If all competitors at or near a highly capable frontier model can make strong containment arguments, Anthropic will meet or exceed their overall risk-reduction posture and, "until we are able to do so, we will delay AI development and deployment as needed." If a competitor implements a superior mitigation at comparable or lower cost, Anthropic will make a significant effort to match it but "will not necessarily delay AI development and deployment in this scenario." Anthropic states these are "necessarily high-level and limited" and that in many cases it "will not have enough information to determine that the relevant scenario applies."
Frontier Safety Roadmap and Risk Reports. A new requirement sets ambitious goals across Security, Alignment, Safeguards, and Policy, shared with all full-time employees, the Board, and the Long-Term Benefit Trust and published in redacted form. Its status is explicitly not that of a commitment — "these are not hard commitments but rather public goals against which we will openly grade our progress" — paired with an anti-downgrade undertaking to "avoid situations where we revise the goals in a less ambitious direction simply because we are unable to achieve them." The stated purpose is organizational: to create a forcing function for work that "can be at cross-purposes with immediate competitive and commercial priorities." This is a partial answer to the public-reporting column in the accountability table above, though the reports are self-produced and redacted rather than externally verified.
Deployment-time governance
The four frameworks compared above locate their decision points before or around release. A distinct approach places the binding constraint during deployment: releasing gradually, monitoring the running system, and pausing or rolling back when behaviour warrants (Iterative Deployment).
OpenAI restated this position on July 20, 2026 after failures observed during limited internal deployment of a long-horizon model, concluding that "no fixed evaluation suite can anticipate every behavior, so pre-deployment testing must be paired with close monitoring, safeguards that can intervene, and the ability to pause or roll back when needed," and arguing that for long-running agents "each step can look acceptable on its own while the sequence can produce an outcome that would not be approved" (Safety and Alignment in an Era of Long-Horizon Models (OpenAI, July 2026)). The safeguards OpenAI describes — incident-derived evaluations, trajectory-level monitoring that can pause a session, and user visibility into long-running sessions — are of a kind none of the four compared frameworks specifies. The limit of the approach is recorded in the companion disclosure published the following day: those deployment safeguards "were intentionally not enabled" during the capability evaluation in which models escaped a sandbox and reached Hugging Face's production infrastructure (OpenAI and Hugging Face Partner to Address Security Incident During Model Evaluation (OpenAI, July 2026)). See Post-Deployment AI System Monitoring.
The Frontier Compliance Framework as bridge to regulation
Anthropic's Frontier Compliance Framework (February 2026) maps the voluntary RSP to binding regulatory requirements (CA SB 53, EU AI Act). It defines a three-tier structure: the RSP as voluntary best practice that goes beyond regulatory requirements; the FCF as a compliance floor reflecting what regulation currently requires; and the EU AI Act and CA SB 53 as binding law. Under this architecture the RSP can evolve as capabilities advance while the FCF translates legal obligations into operational safety practices, providing a worked example of how voluntary frameworks interact with binding regulation.
Anthropic's Advanced AI Framework (June 2026) proposes the binding regime the FCF would map onto, converting several of the voluntary elements above into statutory obligations on "Covered Developers" — a conjunctive test requiring both more than 10²⁵ training FLOP and either more than $500 million in annual AI-derived revenue or more than $1 billion in annual AI R&D spending. Its obligations include a published safety framework, six-monthly risk reports, system cards, 15-day critical-safety-incident reporting, mandatory independent evaluation, a security program, and civil penalties scaling to global annual revenue, across four enumerated risks: biological weapons, offensive cyber operations, loss of control, and automated R&D that could accelerate or amplify the first three (Anthropic's Advanced AI Framework (June 2026)). Its stated position on preemption is restrictive: Congress "should not preempt state law unless it enacts a rigorous federal regime that meets or exceeds the strongest measures proposed in this framework," and compliance with such a regime "should not itself confer immunity, a safe harbor, or a presumption against liability" under state law. Read alongside the RSP v3.4 split described above, the two documents divide the same problem: the RSP states what Anthropic will do unilaterally, and the Advanced AI Framework states what it argues government should require of everyone.
Key tensions
Voluntary vs. mandatory. All lab frameworks are voluntary. Labs argue this allows faster iteration and genuine commitment, while critics argue self-regulation creates a conflict of interest and race-to-the-bottom dynamics (Source: AI Race Dynamics).
Precautionary vs. evidence-based. GPT-5.3 Codex was classified High Cybersecurity on precautionary grounds, whereas RSP thresholds require demonstrating capability levels before triggering restrictions. A more precautionary stance slows deployment; a more evidence-based stance risks deploying before risks are fully characterized.
Frontier vs. broad coverage. The lab frameworks (RSP, Preparedness) focus on catastrophic frontier risks, while the NIST RMF addresses the much larger population of narrower AI systems. The two domains rarely interact, but as agentic AI and GPAI proliferate, the same model can be both a frontier safety concern and an organizational risk management challenge.
Self-assessment vs. external audit. All current frameworks rely primarily on self-assessment. The Safety Cases paper argues for claims that regulators can verify, and building external evaluation capacity (more evaluators like METR, and government AI safety institutes) is a prerequisite for that approach. Both of Anthropic's 2026 documents point the same way from different directions: RSP v3.4 states that "the best way for these recommendations to be implemented is likely via governance of all relevant frontier AI developers by third parties," and the Advanced AI Framework proposes mandatory independent evaluation as a statutory obligation (Anthropic Responsible Scaling Policy v3.4 (July 2026), Anthropic's Advanced AI Framework (June 2026)).
Unilateral safety vs. ecosystem risk. RSP v3.4 introduces a tension the earlier frameworks did not state: whether a developer should hold to an absolute risk standard regardless of what competitors do, or condition its commitments on theirs. Anthropic's answer is to do both in separate registers — unilateral commitments in the main text, more ambitious recommendations it "cannot commit to following… unilaterally," and competitor-contingent delay commitments in Appendix A (Anthropic Responsible Scaling Policy v3.4 (July 2026)). Whether conditioning safety commitments on competitor behaviour mitigates or institutionalizes race dynamics is contested (Source: AI Race Dynamics).
Sources
- Anthropic RSP Version 3.1
- Anthropic RSP Version 3.4 (July 8, 2026)
- OpenAI Preparedness Framework V.2
- NIST AI RMF 1.0
- Safety Cases for Frontier AI (GovAI, 2024)
- Frontier Compliance Framework
- Anthropic's Advanced AI Framework (June 2026)
- Safety and Alignment in an Era of Long-Horizon Models (OpenAI, July 2026)
- OpenAI Hugging Face incident disclosure (July 2026)
- Claude Mythos Preview System Card
- GPT-5.3-Codex System Card
- Taxonomy of Systemic Risks from GPAI
- The Emergence of AI Ethics Auditing
Relationships
- depends-on: Responsible Scaling Policy (RSP), AI Safety Cases and Frameworks — the concepts this comparison operates over.
- related: Iterative Deployment — the deployment-time approach compared under "Deployment-time governance".
- related: Frontier AI Governance, Risk-Based AI Regulation, US AI Regulatory Approaches Compared.
- regulated-by: California SB 53, EU AI Act (Regulation 2024/1689) — the binding law the FCF maps the voluntary frameworks onto.