Version 3.4 of Anthropic's Responsible Scaling Policy took effect July 8, 2026. The RSP is described as "our voluntary framework for managing catastrophic risks from advanced AI systems," establishing "how we identify and evaluate risks, how we make decisions about AI development and deployment, and… how we aim to make sure that the benefits of our models exceed their costs."
The document is the ninth published version: v1.0 (September 19, 2023), v2.0 (October 15, 2024), v2.1 (March 31, 2025), v2.2 (May 14, 2025), v3.0 (February 24, 2026), v3.1 (April 2, 2026), v3.2 (April 29, 2026), v3.3 (May 26, 2026), v3.4 (July 8, 2026). Anthropic states it has "always intended for our RSP to be a living document."
The policy uses "catastrophic risk" in a deliberately non-statutory sense — "risks of the most severe potential harms from advanced AI, such as existential threats or fundamental destabilization of global systems… in its plain meaning rather than adopting any specific statutory definition," noting that where laws such as California SB 53 define the term with specific thresholds, "we address those requirements in separate compliance frameworks."
The structural change: unilateral commitments separated from recommendations
The third iteration's defining move is to split what Anthropic commits to do from what it thinks the industry should do, and the stated reason is a collective action problem.
The prior RSP "committed to implementing mitigations that would reduce our models' absolute risk levels to acceptable levels, without regard to whether other frontier AI developers would do the same." Anthropic now argues this was the wrong unit of analysis: "from a societal perspective, what matters is the risk to the ecosystem as a whole. If one AI developer paused development to implement safety measures while others moved forward with training and deploying AI systems without strong mitigations, that could result in a world that is less safe—the developers with the weakest protections would set the pace, and responsible developers would lose their ability to do safety research and advance the public benefit." It adds a status marker: "Although this situation has not yet arisen, it looks likely enough that we want to prepare for it."
The consequence is stated plainly rather than hedged: the industry-wide recommendations are ones "we cannot commit to following them unilaterally," to be advanced instead "through a mixture of example-setting, addressing unsolved technical problems, advocacy through industry groups, and policy advocacy."
Section 1 presents this as a three-column table: capability thresholds calling for heightened mitigations; Anthropic's own planned mitigations; and the industry-wide recommendation at each threshold.
Why the recommendations are argument-based rather than ASL-based. Anthropic states it "cannot presently give highly specific advance detail on what evaluations will determine whether risk thresholds have been passed, or what risk mitigations will be needed," so the recommendations are "structured around requiring analysis and arguments making a strong case for safety, rather than AI Safety Levels." It names the cost of that flexibility: "one actor's view of what constitutes good risk assessment and mitigation may be very different from another's."
Its stated preferred resolution is third-party governance rather than self-regulation: "the best way for these recommendations to be implemented is likely via governance of all relevant frontier AI developers by third parties that determine which developers need to provide risk analyses… and determine which such arguments are adequate." Where that takes the form of national regulation, "different countries should attempt to harmonize their governance, including standards of evidence, to avoid a race to the bottom"; in the shorter run, standards-setting organizations and auditors "might review such arguments and enforce high quality standards… via voluntary mechanisms." See Frontier AI Governance, Anthropic's Advanced AI Framework (June 2026).
Appendix A — competitor-contingent commitments
Where the main text declines to commit unilaterally, Appendix A commits conditionally. Anthropic frames these as an attempt to "avoid an inadvertent 'race to the bottom' on safety," achievable "to the extent that other relevant AI developers prioritize safety and invest in legible demonstrations that they are doing so."
| Scenario | Commitment |
|---|---|
| Anthropic in the lead — has or will imminently develop a highly capable model, with clear evidence no competitor will soon develop one | Require a strong argument that catastrophic risk is contained, along the lines of the Section 1 recommendations. "We will delay AI development and deployment as needed to achieve this, until and unless we no longer believe we have a significant lead." |
| Competitors have strong safety measures — strong evidence that all competitors at or near a highly capable frontier model can make strong containment arguments | Meet or exceed those competitors' overall risk-reduction posture, as best assessable. "Until we are able to do so, we will delay AI development and deployment as needed." |
| General upleveling — strong evidence a competitor has implemented a mitigation that significantly improves on Anthropic's analogous one and is implementable at comparable or lower cost | Make a significant effort to meet or exceed that standard. "However, we will not necessarily delay AI development and deployment in this scenario." |
Anthropic states two limits on these: they are "necessarily high-level and limited," and "in many cases, we will not have enough information to determine that the relevant scenario applies and will have to use our best judgment." It also states they are not exhaustive of when it might pause: the commitments "do not preclude us from taking cautionary action… we would strongly consider pausing development and/or deployment to improve the safety profiles of our models even in cases not covered below."
Frontier Safety Roadmap
A new requirement in this iteration. The Roadmap lays out "ambitious but achievable goals for improving our risk mitigations" across Security, Alignment, Safeguards, and Policy. It is shared with all full-time employees, the Board, and the Long-Term Benefit Trust, and published in redacted form, with updates on whether goals were achieved and new goals set when they are.
Its status is explicitly not that of a commitment — "these are not hard commitments but rather public goals against which we will openly grade our progress" — and Anthropic states the anti-gaming constraint it is imposing on itself: "we will strive to avoid situations where we revise the goals in a less ambitious direction simply because we are unable to achieve them."
The stated purpose is organizational leverage: to "create a forcing function for work that would otherwise be challenging to appropriately prioritize and resource, as it requires collaboration (and in some cases sacrifices) from multiple parts of the company and can be at cross-purposes with immediate competitive and commercial priorities."
Risk Reports
Also new in this iteration, and distinguished from system cards: Risk Reports "have significant content in common with system cards, but we are adding additional structure and process aimed at presenting our overall assessments of risk," and "unlike system cards, Risk Reports will not be published with each new model release."
Scope and timing. Published every 3–6 months, covering a coverage date no more than 30 days before publication. In scope: all publicly deployed models as of the coverage date, plus internally deployed models determined to pose risks related to high-stakes misalignment or automated R&D that significantly exceed those in a prior Risk Report — "at a minimum… any internal models that we are deploying for large-scale, fully autonomous research."
Off-cycle updates are required in two circumstances: on public deployment of a model significantly more capable than all models whose chemical/biological, high-stakes-misalignment, or automated-R&D risks have been publicly analyzed; and within 30 days of determining that an internally deployed model exceeds all such previously analyzed models on misalignment or automated-R&D risk.
Contents. Threat model identification and specification; evidence including evaluations of relevant capabilities and behaviors; risk mitigations implemented as of the coverage date; and additional material factors. The assessment section covers threat-specific remaining absolute risk, an overall risk assessment, a risk-benefit determination explaining "whether, and if so why, we believe the identified risks are" justified, and forward-looking monitoring plans. Reports also disclose deviations: noteworthy cases where mitigation practices departed from the norm, decisions to internally deploy in-scope models that would not otherwise be reviewed, and "changes to our Frontier Safety Roadmap and any cases where we failed to meet our goals."
Where competitor comparison bears on the assessment, reports cover competitive landscape analysis, the role of those comparisons in the risk assessment, benefits analysis, and "advocacy efforts: the steps we took to raise public awareness of the relevant risks."
Stated principles. Reports are to be "direct, candid, and informative," and specifically: "we will acknowledge when we view certain models as posing significant risks in absolute terms, even if our marginal contribution to overall ecosystem risk may be relatively limited." Marginal risk analysis — arguing that Anthropic's own systems impose relatively lower risk given risks unavoidably posed by others — triggers a governance escalation: where it "plays a major role in a decision to move forward, explicit approval of the Risk Report by the Board and LTBT (rather than just the CEO and RSO) will be required."
Redactions. Four permitted grounds: legal compliance (export control, national security, contractual obligations), intellectual property, public safety, and privacy. Anthropic commits to "disclose the existence of each redaction made in the public version… and aim to give a brief justification," while noting "in many cases this will necessarily be very high-level."
External review. Anthropic commits to "work toward a practice of seeking comprehensive, public external review," under which third-party organizations receive unredacted or minimally redacted versions and publish commentary addressing "the quality of our reasoning, the validity of our risk assessments, the overall level of risk, and whether the redactions we've made for the public version are reasonable and appropriate." The stated aim is that "external reviewers' judgments… carry significant weight in the eyes of the public."
Governance
Following approval of a Risk Report, the CEO and the Responsible Scaling Officer share the decision, the underlying report, and internal feedback with both the Board and the Long-Term Benefit Trust. The marginal-risk escalation above shifts approval authority from CEO and RSO to Board and LTBT.
Relationships
- supersedes: Anthropic's Responsible Scaling Policy (Version 3.1) — later version in the same v3.x series
- related: Anthropic's Responsible Scaling Policy (Version 2.2) — the prior v2.x structure this iteration departs from
- supports: Responsible Scaling Policy (RSP) — the current text of the framework the genre is named after
- related: Anthropic's Advanced AI Framework (June 2026) — the third-party-governance outcome this policy names as preferable is what that framework proposes in legislative form
- related: California SB 53 — addressed through separate compliance frameworks rather than through the RSP's own definition
- related: Anthropic, Frontier AI Governance, AI Safety Cases and Frameworks