AI Policy Wiki
Dashboard

Open Problems in Frontier AI Risk Management

high confidence · updated 2026-06-06

Multi-author (Oxford / MIT / Stanford / Purdue / Vilnius / Mercatus / Saferai / Concordia / Pivotal / CARMA) systematic survey of unresolved challenges in frontier AI risk management across the five-stage ISO 31000 lifecycle. Argues capability thresholds in current frontier-lab safety frameworks (RSP, Preparedness, FSF) measure proxies rather than real-world risk. Published as a Substack-hosted preprint May 4, 2026.

Open Problems in Frontier AI Risk Management is a multi-author working paper that surveys unresolved challenges in managing the risks of frontier general-purpose AI, organized around the five-stage ISO 31000 risk management process. It was published May 4, 2026 and disseminated via Andrew Clearwater's "AI Governance Stack Has Holes" Substack. The paper is framed as a problem-oriented, agenda-setting reference document accompanied by a living online repository, and it does not propose solutions.

Authorship and provenance

The named authors are Marta Ziosi, Miro Plueckebaum, Stephen Casper, Henry Papadatos, Ze Shen Chin, Peter Slattery, James Gealy, Tim G. J. Rudner, Brian Tse, Ariel Gil, Patricia Paskov, Maximilian Negele, Rokas Gipiškis, Nada Madkour, Vera Lummis, Rupal Jain, Luise Eder, Kristina Fort, Malou C. van Draanen Glismann, Inès Belhadj, Amin Oueslati, Anna K. Wisakanto, Richard Mallah, Koen Holtman, Ranj Zuhdi, Daniel S. Schiff, Jessica Newman, Malcolm Murray, and Robert Trager.

Representative affiliations span the Oxford Martin AI Governance Initiative, MIT CSAIL, MIT Future Tech, Stanford, Purdue (Governance and Responsible AI Lab), UC Berkeley CLTC, Mercatus / GMU, Vilnius University, SaferAI, AI Standards Lab, The Future Society, Concordia AI, Pivotal Research, the Center for AI Risk Management & Alignment (CARMA), Vijil, and the University of Toronto. The authorship is broad and cross-institutional, spanning 17 distinct institutions that include industry safety labs (SaferAI, Concordia AI), think tanks (Mercatus, The Future Society, Pivotal), and academic governance centers (Oxford Martin, Berkeley CLTC, Purdue Governance and Responsible AI Lab); the dev-log described the contributing institutions more narrowly as "Oxford, MIT, Stanford, and UC Berkeley." The authors note that inclusion as an author does not entail endorsement of all aspects of the paper.

Core argument

The paper holds that most existing AI-specific risk management standards — ISO/IEC 23894:2023 and ISO/IEC 42001:2023 — were developed for narrow AI systems before the emergence of frontier general-purpose AI. It argues that frontier AI both amplifies existing risks and introduces qualitatively novel challenges, citing a lack of stable scientific consensus owing to the rapid pace of technological change, and arguing that emerging frontier-AI safety practices are often misaligned with, or may undermine, established risk management frameworks.

The paper surfaces open problems across the five stages of the ISO 31000 risk management process — risk planning, risk identification, risk analysis, risk evaluation, and risk mitigation — and classifies each open problem into one of three categories: (a) lack of scientific or technical consensus; (b) misalignment with, or challenges to, established risk management frameworks; and (c) shortcomings in implementation despite apparent consensus and alignment.

Five-stage structure

The survey is organized by the five ISO 31000 stages and their sub-stages:

  1. Risk planning — establishing the scope and context (1.1); setting objectives (1.2); setting criteria (1.3).
  2. Risk identification — identifying risk sources (2.1); identifying potential events, controls and consequences (2.2).
  3. Risk analysis — internal information gathering (3.1); external information gathering (3.2); severity of consequences and likelihood (3.3).
  4. Risk evaluation — determining risk acceptance (4.1); deployment decisions (4.2).
  5. Risk mitigation — data-level mitigations (5.1); model-level mitigations (5.2); system-level mitigations (5.3); ecosystem-level mitigations (5.4).

Key findings

A central argument is that the capability thresholds used by current frontier-lab safety frameworks — Anthropic's RSP, OpenAI's Preparedness Framework, and Google DeepMind's Frontier Safety Framework — measure proxies (for example, a model's score on a cyber-uplift evaluation) rather than real-world risk (the probability and magnitude of actual harm). The paper treats this misalignment with established risk-management practice (ISO 31000) as one of its central concerns.

Across the five stages, the paper catalogues 27 or more unresolved problems. Examples cited in Andrew Clearwater's commentary include: how to set frontier-AI risk acceptance criteria when the harm distribution is unknown and possibly heavy-tailed; how to operationalize "marginal risk" or "uplift" without a defensible counterfactual baseline; in internal-information gathering, developers' incentive to under-report capability findings and questions of auditor independence; in external-information gathering, the tension between model-card and system-card transparency and IP or safety considerations; the observation that ecosystem-level mitigations (such as monitoring and deployment-context controls) are often more tractable than system-level versus model-level mitigations but are rarely the focus of frontier-lab framework documents; and, in severity-of-consequences modeling, the lack of structured methods for AI-specific catastrophic-risk pathways.

The paper deliberately does not propose solutions. It instead provides what it describes as "a problem-oriented, agenda-setting reference document" intended to support coordination, reduce duplication, and guide future research and governance efforts.

Methodology

The approach is a literature review with structured open-problem identification rather than empirical research, grounded in the Grant & Booth (2009) review framework. The document runs to 81 pages and catalogues 27 or more open problems. A living online repository, referenced in the paper, hosts updates as the field evolves.

Reception

The paper was surfaced in dev-log items including the May 4, 2026 Reuters/NYT thread on the Trump administration's reported pre-release vetting executive order (AI Pre-Release Vetting), where it was framed as evidence that even the leading frontier-AI labs' own published safety frameworks would not survive a rigorous risk-management audit. Andrew Clearwater's Substack commentary characterized the paper's contribution as exposing how "the AI governance stack has holes in it." (Source: andrewclearwater.substack.com)

Relation to other work

The paper relates to AI Safety Cases and Frameworks by arguing that capability thresholds are necessary but not sufficient and need to be embedded in a full ISO 31000-style risk-management process to deliver "robust and meaningful consensus." It runs against the implicit framing in some frontier-lab safety-framework documents that capability-threshold-driven precommitments are themselves a complete safety case. It overlaps with the Frontier Model Forum's Managing Advanced Cyber Risks paper, in that both call for tighter alignment between frontier-lab practice and broader risk-management standards. It builds on Risk Taxonomy: Catastrophic, Systemic, and Existential Risk and Risk-Based AI Regulation by applying a structured taxonomy at the procedural level rather than the harm-vector level. The paper cites the UK AI Safety Institute as one of the authoritative producers of "easily updatable, specific technical guidance."

Relationships

Citation

Ziosi, M., Plueckebaum, M., Casper, S., et al. (2026). Open Problems in Frontier AI Risk Management. Working paper. Oxford Martin AI Governance Initiative et al.