AI Policy Wiki
Dashboard

UK AISI / CAISI Preliminary Assessment of Kimi K3's Cyber Capabilities (July 2026)

high confidence · updated 2026-07-25

Joint UK AISI and NIST CAISI preliminary cyber evaluation of Moonshot AI's Kimi K3, finding it significantly below frontier US models (step 17 of 32 on 'The Last Ones' against 28.5; 0 of 41 arbitrary code execution against 20) but above GLM-5.2, with safeguards that did not prevent agentic exploit development.

The preliminary assessment of Kimi K3's cyber capabilities is a joint evaluation by the UK AI Security Institute and the US Center for AI Standards and Innovation, published July 23, 2026 on nist.gov with a mirror on aisi.gov.uk. It is the first government evaluation of Moonshot AI's Kimi K3, released July 16, 2026 and slated for open-weight release by July 27, 2026.

The assessment makes three headline findings: Kimi K3 performs significantly below the most recent frontier cyber-capable models; it performs above GLM-5.2 on the same evaluations; and its safeguards "did not prevent it from attempting cyber exploit development or offensive cyber operations."

Method and caveats

The evaluators describe the results as preliminary, run on "a small set of public and private benchmarks." Two methodological caveats shape the comparison:

  • US models were evaluated with system-level safeguards disabled to reduce refusals and enable measurement of maximal capability. Publicly available versions of those models have the safeguards enabled.
  • Kimi K3's overall cyber-capability estimate carries a wider confidence interval than the other models', because it was derived from a single benchmark (ExploitBench, 41 tasks focused on exploit development), whereas the other models' aggregate scores drew on a larger task set covering additional cyber domains. The restriction is attributed to "the specifics of Kimi K3's hosting setup," which allowed only a selective set of evaluations.

Aggregate cyber capability is computed across tasks from multiple benchmarks using an approach the report describes as inspired by Item Response Theory, with the methodology set out in CAISI's earlier published assessment of Z.ai's GLM-5.2. On the aggregate trend chart, a 400-point increase corresponds to a tenfold increase in the odds of solving tasks.

Exploit development: ExploitBench

ExploitBench is a public benchmark developed by Carnegie Mellon University (arXiv 2605.14153) measuring a model's ability to progress along the software exploitation ladder — coverage and crash reproduction, arbitrary read/write, control-flow hijack, and arbitrary code execution. It tests models against 41 post-2023 vulnerabilities in the V8 engine, the JavaScript and WebAssembly engine used by Chrome.

MeasureKimi K3GLM-5.2Most cyber-capable models
ExploitBench score32%24%
Arbitrary code execution (ACE)0 of 4120 of 41 (average)

The report describes arbitrary code execution as "the highest-severity outcome in exploit development, granting attackers the ability to hijack a target," and notes that unlike the most cyber-capable models, Kimi K3 failed to develop exploits achieving it on any sample.

Cyber range: The Last Ones

"The Last Ones" (TLO) is a 32-step simulated corporate-network attack spanning 4 subnets and roughly 20 hosts, which the evaluators estimate would take a human expert about 20 hours. Runs are capped at 100M tokens.

MeasureKimi K3GLM-5.2Most cyber-capable US models
Average steps reached (of 32)171128.5
Full solves1 of 10 attempts6/10 and 7/10 for the most capable

The single successful run is the assessment's most-qualified finding. The report states it "indicates that Kimi K3 is capable of autonomously attacking small, weakly defended and vulnerable enterprise systems, when directed to do so and given initial network access," while immediately noting that TLO "differs from real-world environments in several ways: it lacks active defenders and defensive tooling, imposes no penalty for actions that would trigger security alerts, and contains an intentional attack path."

The evaluators also record a diffusion point: "Solves of TLO are no longer exclusive to a small set of models." Four publicly released closed-weight models had solved the range in prior testing, the most capable of them reliably; Kimi K3 solved it once in ten attempts within the standard token limit.

Relation to the open-weight cyber gap

The assessment cites AISI's July 17, 2026 open-weight cyber-gap post for its description of GLM-5.2 as "the most cyber-capable open-weight model as of June 2026," making that post the methodological predecessor and GLM-5.2 the open-weight comparator. Read together, the two documents place Kimi K3 above the prior open-weight leader while still below the closed frontier.

Reception

Presidential AI adviser David Sacks cited the below-frontier result against AI guardrail requirements (Source: insideaipolicy.com). The safeguards finding cut the other way in the same debate, arriving during the July 2026 argument over restricting Chinese open-weight models. See Open-Weight Frontier Models.

Provenance

Published on nist.gov, a canonical US government host, with a UK AISI mirror at aisi.gov.uk. Article metadata records publication July 23, 2026. Full text pulled and verified July 25, 2026; verification trail at Wiki/_meta/queue/gap-scan/proposed-sources/aisi-caisi-kimi-k3-cyber-2026.md.

Relationships