AI Policy Wiki
Dashboard

UK AISI Frontier AI Trends Report 2025

high confidence · updated 2026-06-06

Inaugural UK AI Security Institute trends report consolidating two years and 30+ frontier models of evaluation data across cyber, SWE, chem/bio, autonomy, safeguards, and societal impact.

The Frontier AI Trends Report 2025 is the inaugural consolidated empirical retrospective from the UK AI Security Institute (AISI), published by AISI within the Department for Science, Innovation and Technology (DSIT) on 2025-12-18. It draws on two years of AISI evaluations (November 2023 – October 2025) across more than 30 frontier systems. The report is framed as a trend document rather than a benchmark ranking, and names no specific models. Its public URL is https://www.aisi.gov.uk/research/aisi-frontier-ai-trends-report-2025.

Scope and method

The report consolidates evaluations of more than 30 frontier systems run between 2022 and October 2025. Evaluations spanned six modalities: auto-graded task sets, Long-Form Tasks, scaffolded agent tasks, expert red-teaming, human-uplift studies, and human-impact studies, with a default protocol of 10 runs per task per model. The domains covered are cyber, software engineering, chem/bio, autonomy and self-replication, safeguards, and societal impact. Evaluations used the AISI Inspect framework, and models were anonymised under a Red/Purple/Green/Blue/Yellow scheme; AISI does not name the systems it tested.

Findings

The headline measurements are summarised below, comparing a late-2023 or earlier baseline against 2025.

DomainMetricLate 2023 / baseline2025
Cyber — apprenticesolve rate9%~50%
Cyber — expert (10+ yrs exp.)any successful completionnoyes (2025, first)
Cyber — task horizondoubling time~8 months
SWE — hour-long taskssolve rate<5%>40%
Chem/BioPhD baseline (38–48%)at parity (May 2024)surpassed (bio far ahead, chem catching up)
Jailbreak discoveryexpert red-teamer time~10 min (older gen)>7 hours (6 months later) — ~40×
AI companionship (UK)any emotional-support use33% annually, 8% weekly, 4% daily

In offensive cyber, apprentice-level task solve rates rose from 9% in late 2023 to roughly 50% in 2025. AISI also records the first observed instance, in 2025, of a model completing an expert-level cyber task calibrated to 10-plus years of professional experience. The cyber task horizon — the length of task a model can complete — was measured to be doubling roughly every 8 months. In software engineering, the solve rate on hour-long tasks rose from under 5% to over 40% across the two-year window.

On chem/bio, models reached parity with a PhD-expert baseline (scored at 38–48%) around May 2024 and have since surpassed it, with biology capability far ahead and chemistry capability catching up. On safeguards, AISI reports that universal jailbreaks exist for every tested system, but that the expert red-teamer time required to discover one rose from about 10 minutes for an older generation to more than 7 hours for a generation tested six months later, an increase of roughly 40×.

For autonomy and self-replication, the report draws on the RepliBench evaluation, finding several frontier models improving. Models were comparatively strong at early-stage self-replication subtasks (acquiring compute, acquiring money) and weaker at later-stage subtasks (maintaining persistent access, evading detection), a pattern AISI describes as consistent with the general task-horizon picture in Measuring AI Ability to Complete Long Software Tasks.

On societal impact, a UK survey (n=2,028) found that 33% of UK adults had used AI for emotional support in the past year, with 8% doing so weekly and 4% daily.

The adaptation buffer

AISI frames the rising jailbreak-discovery times as an "adaptation buffer": the additional effort required to find a universal jailbreak gives defenders lead time over malicious actors. The report presents this buffer as explicitly partial, stating that "safeguards won't prevent all AI misuse." AISI estimates that open-weight systems trail closed frontier capability by roughly 4–8 months, and treats this lag as the real-world buffer window.

Key claims and confidence

The report's principal claims, with AISI's basis and the confidence each carries:

  1. Apprentice-level cyber solve rates rose from 9% to about 50% between late 2023 and 2025 (high; direct longitudinal measurement across 30+ models).
  2. The first model to complete an expert-level cyber task, calibrated to 10-plus years of professional experience, was observed in 2025 (high; AISI direct testing).
  3. Hour-long SWE tasks went from under 5% to over 40% solve rate (high).
  4. Universal jailbreaks exist for every tested system, but the effort to find them is rising about 40× between consecutive generations (high; expert red-team hours-per-jailbreak data).
  5. Open-weight models trail closed frontier capability by 4–8 months (medium-high; AISI estimate).
  6. 33% of UK adults used AI for emotional support in the past year, and 4% daily (high; n=2,028 survey).
  7. The cyber task horizon is doubling roughly every 8 months (medium; based on fewer years of data than METR's general-task curve).

Relation to other AISI and evaluation work

The report is the UK empirical counterpart to US AI Safety Institute — Vision, Mission, and Strategic Goals: where the US AISI vision is a science and coordination strategy document, the UK trends report compiles two years of testing data. It explicitly continues the programme begun in UK AISI — Advanced AI Evaluations May Update (May 2024), which evaluated 5 models, now extended to more than 30 systems across a full two-year window.

The report partially supersedes capability claims in that May 2024 update. On chem/bio, the earlier "at PhD-expert parity" finding has become "surpassed, with biology far ahead," reinforcing the picture in AI Biosecurity. On safeguards, the May 2024 finding that basic jailbreaks defeated all tested models is reframed: all tested models still have universal jailbreaks, but finding one now takes roughly 40× more expert-hours than six months earlier, adding a quantitative dimension to AI Safety Cases and Frameworks that the May 2024 update did not have.

The cyber and SWE trajectories provide longitudinal data on offensive cyber uplift relevant to the threshold debates in Managing Advanced Cyber Risks in Frontier AI Frameworks. The hour-long SWE solve rate moving from under 5% to over 40% in 24 months parallels Measuring AI Ability to Complete Long Software Tasks and AI Software Progress, and the cyber-horizon doubling parallels the general-horizon doubling reported by METR. The 4–8 month open-weight lag is a precise figure on the open-closed capability gap relevant to Compute Governance.

The report extends and refines prior AISI and METR findings rather than contradicting them; no contradictions with existing evaluation work are identified. AISI's refusal to name the models it tested is principled but makes it difficult to map these trends onto specific company or model pages.

April 2026 companion papers

The full PDF, added to Raw Sources/ on 2026-05-02, lists a 54-author contributor list that overlaps substantially with two April 2026 UK AISI papers extending specific findings from the trends report:

The trends report (December 2025) and the sycophancy and sabotage papers (April 2026) were released in close temporal sequence with substantial author overlap across the three publications.

Provenance

Authored by the UK AI Security Institute (AISI), DSIT, and published 2025-12-18 as a flagship agency report. Evaluations were conducted using the AISI Inspect framework with Red/Purple/Green/Blue/Yellow model anonymisation. The full PDF was added to Raw Sources/ on 2026-05-02.

Relationships