AI Policy Wiki
Dashboard

System Card: Claude Opus 4.5 (Anthropic, November 2025)

high confidence · updated 2026-07-26

Anthropic's system card for Claude Opus 4.5. Judges the model below the AI R&D-4 threshold — none of 18 internal survey participants believed it could fully automate an entry-level remote research role — while noting it 'just barely reached our pre-defined benchmark rule-out thresholds.' On CBRN-4 the rule-out is stated as 'less clear for Claude Opus 4.5 than we would like.' Includes a decontamination methodology section and a changelog correcting an ARC-AGI-1 training-set disclosure.

Published November 2025 for Claude Opus 4.5.

Two rule-outs, both qualified

The card is notable less for its determinations than for how narrowly it states them.

AI R&D-4. The threshold requires "the ability to fully automate the work of an entry-level, remote-only Researcher at Anthropic," which the card is careful to distinguish from a weaker reading: "This is a very high threshold of robust, long-horizon competence, and is not merely a stand-in for 'a model that can do most of the short-horizon tasks that an entry-level researcher can do.'"

The determination is negative, supported by a survey: "None of the 18 internal survey participants — who were themselves some of the most prolific users of the model in Claude Code — believed it could fully automate an entry-level remote-only research or engineering role."

But the card flags its own margin: "It is also noteworthy that the model has just barely reached our pre-defined benchmark rule-out thresholds, rather than greatly exceeded them." The specific predicted failures are behavioural rather than knowledge-based — the model "would fail to problem-solve, investigate, communicate, and collaborate in the way a junior researcher could."

CBRN-4. Also ruled out, and also qualified. Opus 4.5 "performed as well as or slightly better than Claude Opus 4.1 and Claude Sonnet 4.5" across biology tasks, but "most notably… in an expert uplift trial, Claude Opus 4.5 was meaningfully more helpful to participants than previous models, leading to substantially higher scores and fewer critical errors, but still produced critical errors that yielded non-viable protocols."

Anthropic states the implication for the framework rather than only for the model: "we take this as an indicator of general model progress where, like in the case of autonomy, a clear rule-out of the next capability threshold may soon be difficult or impossible under the current regime. In fact, the CBRN-4 rule-out is less clear for Claude Opus 4.5 than we would like."

It also attributes part of that uncertainty to the threat model rather than the measurement: "a large part of our uncertainty about the rule-out is also due to our limited understanding of the necessary components of the threat model," CBRN-4 requiring uplift of "a second-tier state-level bioweapons program." See CBRN Uplift.

Decontamination

The card includes a methodology section on benchmark contamination, stating the problem plainly: "When evaluation benchmarks appear in training data, models can achieve artificially inflated scores by memorizing specific examples rather than demonstrating genuine capabilities. This undermines the validity of our evaluation metrics and makes it difficult to compare performance across model generations and among model providers."

Anthropic describes decontamination as "an important component of responsibly evaluating models, albeit one which is an imperfect science," and reports multiple complementary techniques — beginning with substring removal, scanning the training corpus for exact matches and removing documents containing five or more exact question-answer pairs. See AI Benchmarks and Evaluation.

The ARC-AGI correction

The November 24, 2025 changelog records a disclosure correction: Figure 2.10.A was replaced to show semi-private ARC-AGI-1 scores rather than public test-set scores, because "in the previous version of the system card we incorrectly reported that we only trained on the public training set of ARC-AGI-1. We instead trained on a reshuffled train/test split comprising both the public train and test sets."

This is a self-reported error about training-data composition on a benchmark the card reports — relevant context for reading ARC-AGI comparisons across providers.

Relationships