AI Policy Wiki
Dashboard

Safety Cases (Frontier AI)

medium confidence · updated 2026-06-07

A structured, evidence-based argument that an AI system is acceptably safe to develop or deploy in a given context — imported into frontier-AI governance from aviation, nuclear, and defense safety engineering.

A safety case is a structured, documented argument — supported by evidence — that a system is acceptably safe to operate for a defined purpose in a defined environment. The concept is long-established in high-reliability industries: civil aviation, nuclear power, rail, offshore oil, and defense all require a safety case before a system is certified. Its migration into frontier AI governance reframes the question from "did the developer follow a checklist?" to "can the developer affirmatively demonstrate this model will not cause unacceptable harm?"

Contrast with procedural governance

Most early AI governance is procedural — it specifies steps a developer must take (run these evaluations, publish a model card, adopt a policy). A safety case is outcome-oriented: the burden is on the developer to construct a positive argument that the residual risk is acceptable, and to expose that argument to scrutiny. This shifts the default: absent a credible safety case, deployment is not justified.

Structure of an AI safety case

A frontier-AI safety case typically combines several argument types:

  • Inability — the model lacks the dangerous capability altogether (supported by capability evaluations and red-teaming).
  • Control — even if capable, the model cannot exercise the capability because deployment safeguards, monitoring, or affordance limits prevent it.
  • Trustworthiness — the model is capable and unconstrained but will reliably not misuse the capability (the hardest case to make, and the one most dependent on interpretability and alignment evidence).
  • Deference — reliance on the judgment of a credible external body.

Each argument is only as strong as its evidence base, and the genre is explicit that current evidence — especially for trustworthiness — is weak.

Templates and research program

The concept has moved from proposal to worked artifacts. The paper "Safety Cases: A Scalable Approach to Frontier AI Safety" (Buhl et al., 2025) sets out how and why frontier developers might adopt the practice (Source: researchgate.net). The UK AI Security Institute has published a safety-case template for inability arguments, defining safety cases as "clear, assessable arguments that show that a system is safe in a given context" (Source: aisi.gov.uk), and the Centre for the Governance of AI published a template for a cyber inability argument structured in the Claims-Arguments-Evidence (CAE) framework to make safety arguments coherent and explicit (Source: governance.ai). For the harder control case, a published sketch argues models can be deployed safely because of measures such as monitoring and human auditing of suspicious behavior (Source: lesswrong.com). A March 2026 arXiv paper, "Rethinking the Foundations of Frontier AI Safety Cases," surveys this program — including the UK AISI's cyber-inability and control agendas — and argues for revisiting its foundational assumptions (Source: arxiv.org).

Relationship to other frameworks

Safety cases are increasingly the mechanism inside the if-then commitments of a Responsible Scaling Policy (RSP) and its peer AI Safety Cases and Frameworks: at higher capability tiers, the required safeguard is the production of a reviewed safety case. UK government work and bodies such as UK AI Safety Institute (AI Security Institute) and Apollo Research have promoted safety cases as a near-term, auditable governance artifact.

Relationships