AI Policy Wiki
Dashboard

Our evaluation of OpenAI's GPT-5.5 cyber capabilities (UK AISI, April 2026)

high confidence · updated 2026-07-26

UK AISI cyber evaluation of an early GPT-5.5 checkpoint. Reports 71.4% average pass on Expert-level advanced tasks against 68.6% for Mythos Preview, and GPT-5.5 as the second model to complete 'The Last Ones' 32-step corporate-network simulation end-to-end, in 2 of 10 attempts against Mythos Preview's 3 of 10. Also reports a universal jailbreak of the cyber safeguards found in six hours of expert red-teaming, with the final safeguard configuration left unverified.

Published April 30, 2026 by the UK AI Security Institute. See GPT-5.5 ('Spud').

The question it answers

AISI's April evaluation of Claude Mythos Preview had found "a step up in cyber performance over previous frontier models," making it "the first to complete our corporate network attack simulation end-to-end." The open question was whether that was model-specific or general: "A key question was whether this reflected a breakthrough specific to one model, or part of a broader trend."

The finding is that it is a trend: "Results from an early checkpoint of GPT-5.5 suggest the latter: a second model, from a different developer, now reaches a similar level of performance on our cyber evaluations."

Narrow task suite

The suite comprises 95 narrow cyber tasks across four difficulty tiers in capture-the-flag format, testing "vulnerability research and exploitation" through reverse engineering, web exploitation, and cryptography.

The basic tier is described as exhausted: "Models have fully saturated our basic tasks since at least February 2026."

The advanced suite, "built in collaboration with cybersecurity firms Crystal Peak Security and Irregular," targets "realistic targets and modern mitigations, with a significantly larger and more complex search space." The named skills give the difficulty concretely: "reverse engineering stripped binaries and embedded firmware without source; developing reliable exploits for stack and heap overflows, use-after-frees, and type confusions; recovering keys through padding-oracle, nonce-reuse, and weak-RNG attacks; winning TOCTOU races in privileged code paths; unpacking obfuscated malware; and discovering and weaponising synthetic vulnerabilities planted in real open-source software."

Expert-level results (27 Practitioner and 21 Expert tasks, 50M-token budget), average pass rate with one standard error:

ModelExpert pass rate
GPT-5.571.4% (±8.0%)
Claude Mythos Preview68.6% (±8.7%)
GPT-5.452.4% (±9.8%)
Claude Opus 4.748.6% (±10.0%)

AISI's own hedge on the ranking: "On this measure, GPT-5.5 may be the strongest model we have tested" — the error bars on the top two overlap substantially.

Cyber range: The Last Ones

"The Last Ones" (TLO) is "a 32-step corporate network attack simulation, built with SpecterOps… modelled on the kill chain of an enterprise intrusion and spans four subnets and roughly twenty hosts." The agent "starts on an unprivileged attack box with no credentials and must chain together reconnaissance, credential theft, lateral movement across multiple Active Directory forests, a CI/CD supply-chain pivot, and finally exfiltration of a protected internal database." AISI estimates "a human expert would need around 20 hours to complete the full chain."

GPT-5.5 completed TLO end-to-end in 2 of 10 attempts, at a 100M-token budget per attempt — the second model to do so, against Mythos Preview's 3 of 10.

The safeguards result

AISI separates capability from deployed risk explicitly: "The above tests are capability evaluations carried out in a controlled research setting and do not necessarily reflect what is accessible to an ordinary public user of GPT-5.5. Public deployments include additional safeguards, monitoring, and access controls."

It then tested those safeguards, and the result is the report's most operationally significant finding: "We identified a universal jailbreak that elicited violative content across all malicious cyber queries OpenAI provided, including in multi-turn agentic settings. This attack took six hours of expert red-teaming to develop."

The follow-up is left unresolved: "OpenAI subsequently made several updates to the safeguard stack, though a configuration issue in the version provided meant UK AISI were unable to verify the effectiveness of the final configuration." The published record therefore establishes that the safeguards were universally bypassable at the tested version and does not establish that the fix worked.

AISI's reading of the trend

AISI's stated reading: "GPT-5.5 shows that rapid improvement on cyber tasks may be part of a more general trend. If cyber-offensive skill is emerging as a by" — product of general capability improvement rather than targeted training, the implication is that it will appear across developers without being sought. This is the same distinction Anthropic's Sonnet 5 card draws in noting its cyber skill "likely emerges from its general capabilities, rather than targeted training."

Relationships