Version 3.1 of Anthropic's Responsible Scaling Policy (RSP) is the third major iteration of the company's frontier-safety governance document, effective April 2, 2026. It is published by Anthropic. The defining change in this version is that it separates the company's unilateral commitments from recommendations addressed to the wider industry, and it introduces two new instruments: Frontier Safety Roadmaps and Risk Reports.
Summary
Version 3.1 frames frontier safety as a collective action problem rather than a single-company problem. The document states: "If one AI developer paused while others moved forward, that could result in a world that is less safe." On that basis it distinguishes the commitments Anthropic makes on its own from the practices it recommends industry-wide.
The version continues the AI Safety Level (ASL) scheme used in prior iterations, spanning ASL-1 through ASL-4+ with tiered requirements at each level. It develops a sequence of versions: v1.0 (Sep 2023), v2.0 (Oct 2024), v2.1 (Mar 2025), v2.2 (May 2025), v3.0 (Feb 2026), and v3.1 (Apr 2026).
Position in the version sequence
Version 3.1 is not the operative version. Anthropic published three further point releases in the same series — v3.2 (April 29, 2026), v3.3 (May 26, 2026), and v3.4 (July 8, 2026) — and describes the RSP as "a living document." Version 3.4 has been in force since July 8, 2026 and is the ninth published version.
Several elements introduced or foreshadowed here are elaborated in v3.4 rather than left unchanged. The industry-wide recommendations are restructured in v3.4 "around requiring analysis and arguments making a strong case for safety, rather than AI Safety Levels," a departure from the ASL scheme that v3.1 continues to use. The collective-action reasoning is carried forward in extended form and backed by three competitor-contingent commitments in an Appendix A that v3.1 does not contain, two of which commit Anthropic to delaying development and deployment. The Frontier Safety Roadmap, described here as a requirement for concrete plans across Security, Alignment, Safeguards, and Policy, is given an explicit non-commitment status in v3.4 — "not hard commitments but rather public goals against which we will openly grade our progress" — together with an undertaking not to revise goals downward. The Risk Report instrument acquires a publication cadence (every 3–6 months), off-cycle triggers, four enumerated redaction grounds, and a commitment to comprehensive public external review.
Claims about current Anthropic frontier-safety practice should be sourced to v3.4. This page documents the April 2026 text as published.
Key elements
Two instruments are new in this version:
- Frontier Safety Roadmaps — a requirement for concrete plans across four areas: Security, Alignment, Safeguards, and Policy.
- Risk Reports — detailed safety evaluations subject to external review and public publication, with redactions.
The version's organizing premise is that safety is an ecosystem-wide concern. Separating unilateral commitments from industry-wide recommendations is presented as a response to the collective action problem, on the reasoning that a single developer pausing while others continue could leave the overall environment less safe.
RSP in practice: first Sabotage Risk Report
The Anthropic Sabotage Risk Report: Claude Opus 4.6 (Feb 11, 2026) is the first published operationalization of the RSP's Risk Report requirement, and is tied explicitly to the AI R&D-4 standard. It illustrates the form of evaluation the RSP contemplates and sets a template for future R&D-4+ Risk Reports, giving the RSP its first reviewable artifact. Its components are:
- A four-claim structural argument (prior expectations, alignment assessment, inability to undermine the assessment, and limited opaque reasoning/agentic capability).
- Eight explicit catastrophic-sabotage threat pathways, each with frequency × complexity analysis and both alignment and safeguard mitigations.
- Quantitative sabotage-capability evaluations (SHADE-Arena, steganography, subversion strategy).
- A redacted public version plus an unredacted external review by METR (METR's review, March 2026).
Relationships
- related: RSP v3.4 — the version in force since July 8, 2026; that page carries the
supersedes:link to this one. - related: RSP v2.2 — the prior v2.x structure.
- related: Responsible Scaling Policy (RSP) — the framework genre this document names.
- related: AI Race Dynamics — the collective action framing addresses the "race to the bottom on safety".
- related: AI Safety Frameworks — described as the most detailed publicly available frontier safety framework.
- related: Anthropic — core governance document.
- related: OpenAI Preparedness Framework — comparable OpenAI framework.
- related: Seoul Frontier AI Safety Commitments — operationalizes Seoul Commitments II (threshold definitions with home-government input) and IV (public publication of safety frameworks).