The Claude Sonnet 4.6 system card is Anthropic's technical and safety documentation for its mid-tier model Claude Sonnet 4.6, released February 17, 2026. The card reports capability evaluations, an alignment and safety assessment, and the responsible-scaling determinations behind the release decision. Sonnet 4.6 is described as substantially improved over Sonnet 4.5 and as approaching Opus 4.6 on several evaluations, and is deployed under the ASL-3 Standard.
Capabilities and benchmarks
The card states that Claude Sonnet 4.6 substantially improves over Sonnet 4.5 across coding, agentic tasks, reasoning, multimodal capabilities, computer use, and mathematics, and that in several evaluations it approached or matched Claude Opus 4.6, Anthropic's frontier model at the time of release.
Benchmarks evaluated include SWE-bench (Verified and Multilingual); Terminal-Bench 2.0 / OpenRCA; τ2-bench; BrowseComp; and domain evaluations in finance, cybersecurity, and life sciences. On BrowseComp the card reports a highest single-agent score of 74.01% and a multi-agent score of 82.07%, both reflecting post-publication correction (see Correction history).
Sonnet 4.6 introduces thinking modes and an effort parameter, which allow the model to vary its reasoning depth. These are documented in section 1.1.2 of the system card.
Safety and alignment
Sonnet 4.6 is deployed under the ASL-3 Standard, the same standard applied to Sonnet 4.5. The card reports "low overall levels of misaligned behavior" and states that on some alignment measures Sonnet 4.6 showed "the best degree of alignment we have yet seen in any Claude model." Anthropic describes its alignment assessment as covering a "very wide range of potentially misaligned behaviors" and testing model behavior "in unusual and extreme scenarios." An explicit sabotage risk assessment was included in the release decision process.
The release was evaluated against the Responsible Scaling Policy domains. For autonomy risks, the card concludes ASL-3 is appropriate and no ASL-4 trigger was reached. For CBRN risks and for cyber risks, the card reports that no critical threshold was crossed. The ASL-3 Standard applied is defined in RSP v3.1 (Anthropic's Responsible Scaling Policy (Version 3.1)).
Correction history
The BrowseComp scores were updated on March 6, 2026, after an improved cheating-detection pipeline identified 9 instances of unintended solutions in the single-agent setting and 11 in the multi-agent setting. The corrected scores are 74.01% (single-agent) and 82.07% (multi-agent). The card documents these post-publication corrections.
Provenance
The system card was published by Anthropic (Anthropic) on February 17, 2026, accompanying the release of Claude Sonnet 4.6.
Relationships
- related: Claude Sonnet 4.6 — model page for Claude Sonnet 4.6
- related: Claude Opus 4.6 System Card — companion system card for Opus 4.6
- related: Anthropic's Responsible Scaling Policy (Version 3.1) — RSP v3.1 defines the ASL-3 Standard applied
- related: AI Safety Cases and Frameworks — documents safety framework implementation
- related: Anthropic — author