"How Far Behind the Frontier are Leading Open Weight Models on Cyber?" is a post published July 17, 2026 by the UK AI Security Institute under its Cyber & Autonomous Systems programme. It is AISI's first public analysis of how far leading open-weight models trail the closed cyber frontier. Its central finding is that GLM-5.2 and DeepSeek V4-Pro perform comparably to frontier closed models released four to seven months before them — a narrower gap than the six-to-ten-month lag AISI measured in internal evaluations of open-weight models released between January and September 2025.
Framing
AISI states it has tracked frontier cyber capabilities since 2023 and that on its evaluations the most capable models have consistently been closed weight, with leading open-weight models trailing. The post sets out both sides of the open-weight question in the institute's own terms. Open weights allow private hosting with no data returning to the provider, task-specific adaptation, compute-only running costs, a base that providers cannot change or deprecate, open collaboration, and some safety research that requires weight access — including work by AISI's own Model Transparency team.
Against that, AISI argues the same openness "precludes many of the safety measures that closed model developers can use to detect and disrupt misuse, iterate on safeguards as vulnerabilities emerge, control user access and withdraw models," and that once weights are released those options are lost permanently, creating "a persistent and irreversible risk of misuse" for models with dangerous capabilities.
The stated reason the gap matters is preparation time: it is "a window for cyber defenders with access to the most capable closed systems to take action before today's frontier cyber capabilities might become available without the same safeguards." AISI notes that in April 2026 two closed models, Mythos Preview and GPT-5.5, showed some of the largest jumps in AI cyber capability it had observed since testing began, prompting NCSC warnings.
Method
Two evaluation families are used.
Narrow cyber tasks measure the difficulty of tasks a model can complete across four levels — technical non-expert, apprentice, practitioner, and expert — spanning vulnerability research and exploitation, reverse engineering, web exploitation, and cryptography. Results use a 70-task subset of AISI's full 96-task suite to enable historical comparison: 18 technical non-expert tasks (solvable with a technical background but limited cybersecurity experience), 25 apprentice (roughly 1–3 years' experience), 19 practitioner (3–10 years), and 8 expert (10+ years). Each task is given 5 attempts with a 2.5M-token limit per attempt; AISI reports comparator models were unchanged when evaluating at 50M tokens per task against the full 96-task suite.
Cyber ranges measure autonomous end-to-end capability: expert-built simulated networks of hosts, services, and vulnerabilities arranged into sequential attack chains beginning at initial network access. AISI states the ranges currently lack security features of well-defended real environments — active defenders, defensive tooling, and alert penalties. The range reported here is "The Last Ones" (TLO), a 32-step corporate-network attack spanning 4 subnets and roughly 20 hosts, which AISI estimates would take a human expert about 20 hours. Runs are capped at 100M tokens, averaged over 10 runs.
Two open-weight models were selected as candidates to lead open-weight cyber capability at their release: GLM-5.2 (June 2026) and DeepSeek V4-Pro. AISI states it intends to test Kimi K3 on the same basis once its weights are released, announced for end of July.
Findings
| Model | Narrow cyber tasks — comparable closed model | Cyber ranges (TLO) — comparable closed model | |||
|---|---|---|---|---|---|
| GLM-5.2 (Jun 2026) | [[models/claude-opus-46 | Opus 4.6]] and [[models/gpt-53-codex | GPT-5.3-Codex]], released 4 months earlier; holds across all four difficulty levels | [[models/claude-opus-45 | Opus 4.5]], released less than 7 months earlier |
| DeepSeek V4-Pro | [[models/claude-opus-45 | Opus 4.5]], released 5 months earlier | Below [[models/claude-sonnet-45 | Sonnet 4.5]], a sub-cyber-frontier model released 7 months earlier |
On the ranges, AISI notes GLM-5.2 reached step 7 with marginally fewer tokens than any other model on average, tracking Opus 4.6's trajectory to step 11 before stalling. It describes the range-derived gap as larger than the narrow-task gap but treats it as weaker evidence, drawn from a smaller set of ranges than the narrow-task suite, and cautions that trajectories alone do not distinguish stalls caused by insufficient cyber capability from stalls caused by insufficient agentic capability to sustain long-horizon planning.
AISI also states the 4-to-7-month result was found "despite rapid advancements AISI observed in frontier cyber capabilities up to February 2026," and is "not predictive of whether future open weight models will replicate the more recent jumps delivered by Mythos Preview and GPT-5.5."
Safeguards and cost
On safeguards, AISI reports that its evaluations of the tested open-weight models "were largely unimpeded." DeepSeek V4-Pro occasionally refused narrow cyber tasks, mainly in reverse engineering, which AISI circumvented "simply via a small number of repeat attempts at refused tasks." The post argues that deployment-time measures — monitoring, classifiers, user-banning — require control over model access and cannot be applied universally once weights are public, and that remaining techniques such as refusal training "are often easily reversible with access to the weights."
On cost, at advertised first-party prices:
| Measure | Opus 4.5 | Opus 4.6 | GLM-5.2 | DeepSeek V4-Pro |
|---|---|---|---|---|
| 100M-token cyber-range run | ~$85 | ~$85 | ~$46 (est.) | $1.19 |
| Per task solved with 100% reliability | $12.50 | $15.17 | $6.12 | $0.28 |
The per-task figures are pairwise: Opus 4.6 against GLM-5.2, and Opus 4.5 against DeepSeek V4-Pro, across tasks both models in each pair solved with full reliability. AISI notes it did not use first-party providers for the open-weight models tested, so real compute costs may vary.
Stated limitations
AISI states its setup "likely slightly underestimates open weight models' maximum capability," because it did not pursue specific elicitation or optimisations that could have improved performance, and that the post covers cyber capabilities only, so inferences to other capabilities cannot be drawn.
Conclusion and reception
The post concludes that the four-to-seven-month lag "implies cyber defenders have a short window to prepare," endorsing the National Cyber Security Centre's existing encouragement to invest in cybersecurity baselines and AI-enhanced defences. It states that no current approach guarantees safety but that several strategies across model training, safeguarding, auditing, and access could meaningfully reduce risk in combination, and that AISI will continue evaluating leading open-weight models.
The finding became a reference point in the US open-weight policy debate within a week of publication, cited in coverage of the July 2026 Chinese open-weight release wave and in the argument over restricting Chinese models. The joint UK AISI–CAISI assessment of Kimi K3 published July 23, 2026 cites this post as the source for its description of GLM-5.2 as "the most cyber-capable open-weight model as of June 2026." See Open-Weight Frontier Models, US-China AI Competition: Different Races, Different Metrics.
Provenance
Published on aisi.gov.uk, the institute's own domain. The post is dated July 17, 2026; a July 24, 2026 developments-log entry had carried it under a July 24 date, and the canonical publication date is used here. Full text pulled and verified July 25, 2026; verification trail at Wiki/_meta/queue/gap-scan/proposed-sources/aisi-open-weight-cyber-gap-2026.md.
Relationships
- supports: Open-Weight Frontier Models — supplies the measured open-versus-closed gap the policy debate turns on.
- depends-on: UK AI Safety Institute (AI Security Institute) — the evaluating body.
- related: GLM-5.2, DeepSeek V4 Pro / V4 Flash, Kimi K3 — the models evaluated or slated for evaluation.
- related: AI and Cybersecurity, Autonomous cyber-agents, AI Benchmarks and Evaluation — the evaluation domain.
- related: Inference Economics and Token Pricing — the cost-performance comparison.