AI Policy Wiki
Dashboard

Responsible Scaling Policy (RSP)

medium confidence · updated 2026-07-25

Anthropic's if-then governance framework tying model capability thresholds to escalating safety and security commitments — the template for the frontier-lab safety-framework genre.

A Responsible Scaling Policy is a published, pre-commitment governance framework that ties an AI developer's escalation of model capabilities to a corresponding escalation of safety and security measures. Its defining structure is if-then: if a model crosses a defined capability threshold, then a defined set of protective commitments must already be in place before that model is trained or deployed. The term originates with Anthropic, whose RSP was first published in September 2023 and became the template for the broader frontier safety framework genre.

Background and origin

Anthropic published the first version of its Responsible Scaling Policy in September 2023, introducing the AI Safety Level scale and the if-then commitment structure. The framework was subsequently revised, with later versions restructured around capability thresholds and "required safeguards" (Anthropic's Responsible Scaling Policy (Version 2.2)) and then refining thresholds, evaluation protocols, and governance and accountability provisions (Anthropic's Responsible Scaling Policy (Version 3.1)). Three version families are documented: v1.0 (September 2023), which introduced the AI Safety Levels and the if-then structure; the v2.x series; and the v3.x series. Within the v3.x series, version 3.4 took effect on July 8, 2026; it revised the automated AI R&D capability threshold, loosened distribution of unredacted Risk Reports to at least 200 employees, allowed Risk Reports to state coverage dates, required public redaction markers on published materials, and permitted external review to be split across reviewers (Anthropic Responsible Scaling Policy v3.4 (July 2026)).

The v3.x series makes a structural change that distinguishes it from the earlier if-then framing. Anthropic separates commitments it will meet unilaterally from more ambitious industry-wide recommendations it states it "cannot commit to following… unilaterally," on the reasoning that the prior policy had targeted the company's own absolute risk level "without regard to whether other frontier AI developers would do the same," whereas "from a societal perspective, what matters is the risk to the ecosystem as a whole" — a pause by one developer while others proceed could leave "the developers with the weakest protections… set[ting] the pace." The industry recommendations are structured around requiring safety arguments rather than around AI Safety Levels, and Anthropic names third-party governance determining which arguments are adequate as the likely implementation route, with international harmonization of evidentiary standards "to avoid a race to the bottom."

Three competitor-contingent commitments in Appendix A bridge the gap: where Anthropic has a clear lead it will require a strong containment argument and delay development and deployment as needed; where all relevant competitors can make strong containment arguments it will meet or exceed their risk-reduction posture, again delaying as needed; and where a competitor has implemented a materially better mitigation at comparable cost it will make a significant effort to match it, but "will not necessarily delay AI development and deployment in this scenario." The version also introduces Frontier Safety Roadmaps — public goals across Security, Alignment, Safeguards, and Policy, explicitly "not hard commitments," with a stated undertaking not to revise goals downward merely because they proved unachievable — and Risk Reports published every 3–6 months with disclosed redactions, comprehensive external review, and a governance escalation requiring Board and Long-Term Benefit Trust approval whenever marginal-risk reasoning is material to a decision to proceed (Anthropic Responsible Scaling Policy v3.4 (July 2026)).

Key elements

The framework rests on a tiered capability scale, evaluable thresholds, paired families of safeguards, and a commitment to pause when safeguards lag capability.

AI Safety Levels (ASLs) form a tiered scale loosely modeled on biosafety levels (BSL). ASL-2 covers current frontier models; ASL-3 and ASL-4 define successively more dangerous capability regimes. Each level carries a fixed package of required safeguards.

Capability thresholds are the concrete, evaluable triggers that activate the next ASL's commitments. Examples include a model meaningfully uplifting a non-expert's ability to cause a mass-casualty biological or chemical event, or crossing autonomy or cyber thresholds.

Required safeguards fall into two families. Deployment safeguards cover misuse prevention, harm refusal, and monitoring. Security safeguards cover protecting model weights against theft by increasingly capable threat actors.

Evaluations serve as the trigger mechanism: the policy depends on the capability evaluations that decide whether a threshold has been crossed, which makes dangerous-capability evaluation the load-bearing technical input. Under the pause commitment, if a model crosses a threshold and the corresponding safeguards are not ready, the developer commits to pause scaling or deployment until they are.

Relation to policy and the broader genre

The RSP is the first instance of what has become a category of frontier safety frameworks. Peer frameworks adopt the same if-then logic under different names, including OpenAI's Preparedness Framework and Google DeepMind's Frontier Safety Framework; collectively these are the subject of AI Safety Cases and Frameworks. The approach also informs voluntary government regimes, including the commitments brokered at the UK and Seoul AI Safety Summits and the disclosure expectations behind Executive Order 14110 (RESCINDED).

Debates and open questions

RSPs are voluntary and self-enforced. Critics argue they substitute corporate discretion for binding oversight and can be weakened or reinterpreted under competitive pressure. A second line of criticism concerns threshold vagueness: capability thresholds are difficult to operationalize, and an evaluation regime that under-detects dangerous capability defeats the policy.

The framework's relationship to Safety Cases (Frontier AI) is one of convergence. Higher ASLs increasingly require an affirmative safety case — a structured argument that a model is safe to deploy — rather than checklist compliance, bringing the two concepts together.

Relationships