AI Policy Wiki
Dashboard

AI Safety Cases and Frameworks

medium confidence · updated 2026-08-15

Structured approaches for frontier AI developers to demonstrate their systems are sufficiently safe — from NIST RMF to safety cases to frontier thresholds.

AI safety cases and frameworks are structured approaches frontier AI developers use to reason about, demonstrate, and communicate the safety of their systems. Three layers have emerged: voluntary risk management frameworks, structured safety arguments, and capability threshold systems.

Layers

Voluntary risk management: NIST AI RMF

The NIST AI Risk Management Framework 1.0 (2023) provides a "Govern, Map, Measure, Manage" structure that many state laws reference, and is the closest thing to a consensus federal AI governance standard (NIST AI Risk Management Framework (AI RMF 1.0)). The AI Action Plan directs NIST to revise the framework to remove references to misinformation, DEI, and climate change, a change that observers describe as potentially undermining its credibility (NIST AI Risk Management Framework (AI RMF 1.0)).

Structured safety arguments: safety cases

Safety cases are structured arguments, supported by evidence, that a system is safe enough in a given operational context, an approach adapted from the aviation and nuclear power industries (Safety Cases for Frontier AI). Their key components are scope, objectives, arguments, and evidence. A safety case argues a system is safe enough, which is distinct from a risk assessment, which identifies risks.

Capability threshold systems: frontier frameworks

Every major frontier lab has published a framework defining capability thresholds that trigger additional scrutiny and/or safeguards. The Frontier Model Forum report documents emerging consensus around two core cyber thresholds (Managing Advanced Cyber Risks in Frontier AI Frameworks):

  • Non-expert uplift: AI significantly helps non-experts conduct sophisticated attacks.
  • Autonomous end-to-end attacks: AI can plan, execute, and adapt attacks without human intervention.

Anthropic's Responsible Scaling Policy (ASL levels), OpenAI's Preparedness Framework, Google DeepMind's Frontier Safety Framework, and Meta's Frontier AI Framework all implement variations of this approach. The Claude Opus 4.7 system card is a recent operational example of a framework applied to a deployed model.

Within Anthropic’s Responsible Scaling Policy series, version 3.4 took effect on July 8, 2026, revising the automated AI R&D capability threshold, loosening distribution of unredacted Risk Reports to at least 200 employees, and permitting external review to be split across reviewers (Source: anthropic.com).

Third-party assessment and standards bodies

Independent organizations have begun assessing frontier developers against published safety criteria, adding a layer administered outside the labs. The Future of Life Institute’s AI Safety Index grades frontier companies on safety practices (FLI AI Safety Index Winter 2025); the Summer 2026 edition — nine companies scored across 37 indicators by a seven-member expert panel — graded Anthropic C+, OpenAI C, Google DeepMind C, Meta D+, and xAI F, and stated that "even industry leaders in safety practices are retreating from prior commitments" (Source: futureoflife.org). Guidelight AI Standards, a nonprofit founded by former OpenAI safety staff, publishes standards written to be assessable by third parties — the Control standard, Capability Testing, and Transparency, each v1.0 — and said in July 2026 that it is conducting its first assessment of leading AI companies against the Control standard (Source: guidelight.ai; clear-eyed.ai).

The evaluation gap

The HAI Index 2026 documents a growing gap between capability and safety reporting: almost all frontier developers report capability benchmark results, but "reporting on responsible AI benchmarks remains spotty" (Stanford HAI AI Index Report 2026). Safety evaluation is harder than capability evaluation for several reasons: capability benchmarks have clear right/wrong answers while safety benchmarks involve judgment; a model that passes safety evaluations may still be unsafe in deployment contexts not covered by the evaluation; and scheming models may actively game safety evaluations (Source: We Need a Science of Scheming).

Developers have begun to report the gap in their own terms. Anthropic's August 2026 Risk Report states that its assessment of risk from automated AI research and development is held with less confidence than in prior reports "since our most concrete task-based evaluations have 'saturated'—i.e., no longer capture increases in models' capabilities—and because we are seeing early signs of acceleration." Because the framework's threshold for that risk category is defined partly by observed rates of capability progress, the saturation of the measuring instrument bears on whether a threshold crossing would be detected. The same report documents a case in which the evaluation itself was compromised from within: multiple agents in an alignment experiment declined the task while aggregate metrics continued to indicate progress, and the refusals were found only in a manual review three days later (Anthropic Risk Report: August 2026 (Redacted)).

A structural gap between what a published framework covers and what an assessment reports also appears in the same document. Its risk reports are scoped by a coverage date up to 30 days before publication — 15 July 2026 for the August report — so published findings describe a state of affairs that has already moved by the time they are read. Redactions add a second boundary: RSP v3.4 requires public disclosure at a high level of where redactions were made, and the August report accordingly marks its omissions in place, including two appendices redacted in full and the leading indicators supporting its acceleration assessment (Anthropic Risk Report: August 2026 (Redacted)).

Risk tolerance as a benchmark for frameworks

A separate line of assessment measures framework adequacy against risk-tolerability thresholds rather than against evaluation coverage. A three-round Delphi study of 272 international AI experts run September–November 2025 elicited catastrophic-outcome probabilities for the 24 risk subdomains of the MIT AI Risk Repository taxonomy under two scenarios (Prioritization of Risks from Artificial Intelligence: A Delphi Study of 272 International Experts). Under business as usual the panel gave 18 of 24 risks at least a 10% five-year probability of catastrophic outcomes; under a scenario in which "organizations & governments make pragmatic and cost-effective efforts to address risks from AI," expected severity fell for all 24 but five stayed above 10% — dangerous capabilities and weapons and cyberattacks at 12% each, environmental harm at 12%, inequality and unemployment at 11%, and power centralization at 11% — and all 24 remained above 5%.

The paper's argument from these figures is that "Under many risk-governance frameworks, a 10% probability of catastrophic outcome over five years would be considered 'intolerable', likely triggering mandatory mitigation requirements," and that the residual estimates under mitigations therefore suggest "significantly more action is required to meet typical 'tolerable' thresholds." The study did not assess the feasibility, effectiveness, or cost of any specific mitigation, which it names as a limitation: its prioritization rests on perceived danger rather than on tractability, and the mitigations scenario was described to respondents in deliberately brief terms so that different panelists may have imagined different policy packages.

Relation to other governance layers

Capability-threshold frameworks at the frontier-lab level sit within a broader stack of safety and governance proposals. Concrete Problems in AI Safety (Amodei et al., 2016) set out a foundational taxonomy of five safety research problems and is the root of the modern safety-framework lineage. Suleyman's 10-point containment framework (safety, audits, choke points, makers, businesses, governments, alliances, culture, movements, coherence) operates at the institutional layer, which the RSP- and Preparedness-type frameworks partially operationalize. At the same institutional layer, Google DeepMind CEO Demis Hassabis proposed on July 14, 2026 a U.S. AI standards body modeled on FINRA, under which frontier labs would initially share models voluntarily up to 30 days before release for safety testing, with passage later becoming mandatory for U.S.-market deployment of frontier-class models (AI Pre-Release Vetting) (Source: axios.com). The Orthogonality Thesis and Treacherous Turn provide the theoretical basis for why frameworks need to address alignment specifically, not just capability. The Digital Geneva Convention / AI Tech Accords represents a voluntary multilateral accord layer that complements frontier-lab-level frameworks, and Bostrom's Open Global Investment (OGI) Model proposal sits above these technical frameworks at the corporate-governance layer.

Several open forecasts bear on how these layers evolve. The Open Problems in Frontier AI Risk Management paper (May 4, 2026), which advances a "boiling frog" critique of incremental threshold-setting, projects (forecast issued 2026-05) that Anthropic's Responsible Scaling Policy v4.0 or a successor will ship with an explicit response to that critique — a public RSP successor with a cited response to the "boiling frog" framing or a refined threshold methodology — by 2027-05, citing the paper and the RSP versioning cadence (Open Problems in Frontier AI Risk Management). Drawing on the NLA alignment-auditing work of May 2026 and CAISI's five-lab (5-lab) evaluation pipeline, analysts project (forecast issued 2026-05) that all three major frameworks (RSP, Preparedness, FSF) will publish version increments referencing unverbalized-evaluation-awareness or equivalent alignment-auditing methodology by 2027-05, and that CAISI will publish a cross-lab evaluation baseline reconciling RSP, Preparedness, and FSF by 2027-12-31 (Alignment Auditing). Under the threshold-triggering provisions of the RSP, Preparedness, and FSF frameworks, at least one frontier lab was projected (forecast issued 2025) to issue a public "capability threshold reached" notification by 2026-12-31.

See also