AI Policy Wiki
Dashboard

METR Frontier Risk Report, February–March 2026: A pilot assessment of rogue deployment risk at frontier AI companies

high confidence · updated 2026-07-26

METR's entity-based pilot assessment of misalignment risk from AI agents used internally at Anthropic, Google, Meta, and OpenAI, with access to each participant's most capable internal models including raw chains of thought. Organizes evidence into means, motive, and opportunity, and concludes internal agents 'plausibly had the means, motive, and opportunity to start small rogue deployments, but they did not have the means to make them highly robust.' States four specific limits on its own accountability value.

Published May 19, 2026 by METR, covering an assessment window of February 16 to March 16, 2026, with participation from Anthropic, Google, Meta, and OpenAI.

What makes this exercise unusual

The assessment is entity-based rather than model-specific: it evaluates the risk posed by AI agents used inside frontier developers, not the risk of a released model, and is "designed to be repeated periodically rather than tied to public releases."

Each participant provided "access to their most capable internal model(s) at the time of assessment, including raw chains of thought," plus "a wide range of non-public information about the capabilities of the shared model(s), how AI was used and monitored internally, and trends in the pace of progress."

The publication process: METR prepared private reports per participant, participants approved what non-public information could be disclosed, and METR wrote the public report. Critically, "our pilot agreements with participating companies did not give them the right to approve this industry-level publication," which METR calls "a meaningful improvement over previous arrangements with companies that required us to submit certain evaluation reports for review and approval."

The finding

Evidence is organized into means (what harmful actions agents could take), motive (whether they might attempt them), and opportunity (whether attempts could succeed given safeguards).

The assessment target is a "rogue deployment" — "a set of agents running autonomously without human knowledge or permission" — and whether it could be made "robust against varying degrees of security and monitoring measures."

"Overall, we believe that internal agents at the time of our assessment plausibly had the means, motive, and opportunity to start small rogue deployments, but they did not have the means to make them highly robust."

The qualifier carries the conclusion: the constraint identified is persistence and resilience, not initiation. And METR expects it to erode: "Given rapidly advancing capabilities, we expect the plausible robustness of rogue deployments to increase substantially in the coming months." A similar process is tentatively planned for late 2026.

The limits METR states on itself

The report devotes explicit space to why it should not be read as accountability, and the disclosure is more useful than the finding for judging what third-party assessment currently delivers. "This pilot was not designed to provide robust accountability":

  • Silent exit. "We gave participants the option to exit silently from the pilot at any point before approving any non-public information from them to include in the report. This means that any company could have withdrawn partway through the process for any reason." A company that did not like where the assessment was heading could leave without that fact appearing anywhere.
  • Redaction. Participants could "redact or anonymize any non-public information pertaining to them," noted explicitly only where METR judged the redactions "highly relevant to our conclusions about risk, which was a subjective judgment call." METR states plainly that "there were a number of interesting pieces of color supporting our key claims that were removed which did not result in a 'qualified' redaction summary statement or an explicit note elsewhere in the text."
  • Courtesy draft. A non-final draft went to participants a week before publication, "explicitly as a notice rather than a request for approval," during which METR "did make minor edits for accuracy or to add additional information based on participant comments."

The report's redaction summary states that, except where noted, "there was no additional redacted information that was important to our conclusions from any of the participating companies."

Its recommendation follows from the exercise rather than the result: "we believe that periodic third-party assessment of risks from developers' internal use of AI should be adopted throughout the industry."

Sources of evidence

Beyond evaluations on shared internal models and public models, the report draws on information shared by participants, system cards, public literature, and "findings from a recent embedded red-teaming exercise" — the latter initiated separately by an Anthropic staff member, with findings discussed for the first time in this report. METR states it is "excited to explore similar collaborations with other companies."

Relationships