The Frontier AI Trends Report 2025 is the inaugural consolidated empirical retrospective from the UK AI Security Institute (AISI), published by AISI within the Department for Science, Innovation and Technology (DSIT) on 2025-12-18. It draws on two years of AISI evaluations (November 2023 – October 2025) across more than 30 frontier systems. The report is framed as a trend document rather than a benchmark ranking, and names no specific models. Its public URL is https://www.aisi.gov.uk/research/aisi-frontier-ai-trends-report-2025.
Scope and method
The report consolidates evaluations of more than 30 frontier systems run between 2022 and October 2025. Evaluations spanned six modalities: auto-graded task sets, Long-Form Tasks, scaffolded agent tasks, expert red-teaming, human-uplift studies, and human-impact studies, with a default protocol of 10 runs per task per model. The domains covered are cyber, software engineering, chem/bio, autonomy and self-replication, safeguards, and societal impact. Evaluations used the AISI Inspect framework, and models were anonymised under a Red/Purple/Green/Blue/Yellow scheme; AISI does not name the systems it tested.
Findings
The headline measurements are summarised below, comparing a late-2023 or earlier baseline against 2025.
| Domain | Metric | Late 2023 / baseline | 2025 |
|---|---|---|---|
| Cyber — apprentice | solve rate | 9% | ~50% |
| Cyber — expert (10+ yrs exp.) | any successful completion | no | yes (2025, first) |
| Cyber — task horizon | doubling time | — | ~8 months |
| SWE — hour-long tasks | solve rate | <5% | >40% |
| Chem/Bio | PhD baseline (38–48%) | at parity (May 2024) | surpassed (bio far ahead, chem catching up) |
| Jailbreak discovery | expert red-teamer time | ~10 min (older gen) | >7 hours (6 months later) — ~40× |
| AI companionship (UK) | any emotional-support use | — | 33% annually, 8% weekly, 4% daily |
In offensive cyber, apprentice-level task solve rates rose from 9% in late 2023 to roughly 50% in 2025. AISI also records the first observed instance, in 2025, of a model completing an expert-level cyber task calibrated to 10-plus years of professional experience. The cyber task horizon — the length of task a model can complete — was measured to be doubling roughly every 8 months. In software engineering, the solve rate on hour-long tasks rose from under 5% to over 40% across the two-year window.
On chem/bio, models reached parity with a PhD-expert baseline (scored at 38–48%) around May 2024 and have since surpassed it, with biology capability far ahead and chemistry capability catching up. On safeguards, AISI reports that universal jailbreaks exist for every tested system, but that the expert red-teamer time required to discover one rose from about 10 minutes for an older generation to more than 7 hours for a generation tested six months later, an increase of roughly 40×.
For autonomy and self-replication, the report draws on the RepliBench evaluation, finding several frontier models improving. Models were comparatively strong at early-stage self-replication subtasks (acquiring compute, acquiring money) and weaker at later-stage subtasks (maintaining persistent access, evading detection), a pattern AISI describes as consistent with the general task-horizon picture in Measuring AI Ability to Complete Long Software Tasks.
On societal impact, a UK survey (n=2,028) found that 33% of UK adults had used AI for emotional support in the past year, with 8% doing so weekly and 4% daily.
The adaptation buffer
AISI frames the rising jailbreak-discovery times as an "adaptation buffer": the additional effort required to find a universal jailbreak gives defenders lead time over malicious actors. The report presents this buffer as explicitly partial, stating that "safeguards won't prevent all AI misuse." AISI estimates that open-weight systems trail closed frontier capability by roughly 4–8 months, and treats this lag as the real-world buffer window.
Key claims and confidence
The report's principal claims, with AISI's basis and the confidence each carries:
- Apprentice-level cyber solve rates rose from 9% to about 50% between late 2023 and 2025 (high; direct longitudinal measurement across 30+ models).
- The first model to complete an expert-level cyber task, calibrated to 10-plus years of professional experience, was observed in 2025 (high; AISI direct testing).
- Hour-long SWE tasks went from under 5% to over 40% solve rate (high).
- Universal jailbreaks exist for every tested system, but the effort to find them is rising about 40× between consecutive generations (high; expert red-team hours-per-jailbreak data).
- Open-weight models trail closed frontier capability by 4–8 months (medium-high; AISI estimate).
- 33% of UK adults used AI for emotional support in the past year, and 4% daily (high; n=2,028 survey).
- The cyber task horizon is doubling roughly every 8 months (medium; based on fewer years of data than METR's general-task curve).
Relation to other AISI and evaluation work
The report is the UK empirical counterpart to US AI Safety Institute — Vision, Mission, and Strategic Goals: where the US AISI vision is a science and coordination strategy document, the UK trends report compiles two years of testing data. It explicitly continues the programme begun in UK AISI — Advanced AI Evaluations May Update (May 2024), which evaluated 5 models, now extended to more than 30 systems across a full two-year window.
The report partially supersedes capability claims in that May 2024 update. On chem/bio, the earlier "at PhD-expert parity" finding has become "surpassed, with biology far ahead," reinforcing the picture in AI Biosecurity. On safeguards, the May 2024 finding that basic jailbreaks defeated all tested models is reframed: all tested models still have universal jailbreaks, but finding one now takes roughly 40× more expert-hours than six months earlier, adding a quantitative dimension to AI Safety Cases and Frameworks that the May 2024 update did not have.
The cyber and SWE trajectories provide longitudinal data on offensive cyber uplift relevant to the threshold debates in Managing Advanced Cyber Risks in Frontier AI Frameworks. The hour-long SWE solve rate moving from under 5% to over 40% in 24 months parallels Measuring AI Ability to Complete Long Software Tasks and AI Software Progress, and the cyber-horizon doubling parallels the general-horizon doubling reported by METR. The 4–8 month open-weight lag is a precise figure on the open-closed capability gap relevant to Compute Governance.
The report extends and refines prior AISI and METR findings rather than contradicting them; no contradictions with existing evaluation work are identified. AISI's refusal to name the models it tested is principled but makes it difficult to map these trends onto specific company or model pages.
April 2026 companion papers
The full PDF, added to Raw Sources/ on 2026-05-02, lists a 54-author contributor list that overlaps substantially with two April 2026 UK AISI papers extending specific findings from the trends report:
- Ask don't tell: Reducing sycophancy in large language models — Dubois, Ududec, Summerfield, Luettgau (UK AISI) — Dubois, Ududec, Summerfield, Luettgau (April 2026), a controlled-experiment paper on sycophancy mitigation through input-framing reframing, building on the trends report's AI-companionship findings about user emotional dependence.
- Evaluating whether AI models would sabotage AI safety research — Kirk, Souly, Fronsdal, D'Cruz, Davies (UK AISI) — Kirk, Souly, Fronsdal, D'Cruz, Davies (April 2026), a sabotage-propensity evaluation across Mythos Preview, Opus 4.7 Preview, Opus 4.6, and Sonnet 4.6, extending the trends report's autonomy and self-replication evaluations into a more specific safety-research-sabotage threat model.
The trends report (December 2025) and the sycophancy and sabotage papers (April 2026) were released in close temporal sequence with substantial author overlap across the three publications.
Provenance
Authored by the UK AI Security Institute (AISI), DSIT, and published 2025-12-18 as a flagship agency report. Evaluations were conducted using the AISI Inspect framework with Red/Purple/Green/Blue/Yellow model anonymisation. The full PDF was added to Raw Sources/ on 2026-05-02.
Relationships
- supersedes: (partially) the capability claims in UK AISI — Advanced AI Evaluations May Update — particularly on chem/bio parity and long-horizon agent tasks, which have both moved forward.
- supports: AI Benchmarks and Evaluation (institutional empirical series)
- supports: AI Biosecurity (chem/bio parity surpassed)
- supports: AI Safety Cases and Frameworks (adaptation-buffer framing)
- supports: Managing Advanced Cyber Risks in Frontier AI Frameworks (cyber threshold calibration data)
- supports: Measuring AI Ability to Complete Long Software Tasks (cyber-horizon doubling parallels general-horizon doubling)
- related: US AI Safety Institute — Vision, Mission, and Strategic Goals — sibling-institution counterpart; UK empirical to US strategic
- related: UK AI Safety Institute (AI Security Institute) (publishing entity)
- related: AI Scheming, Alignment Faking (self-replication harness adjacent)
- related: Ask don't tell: Reducing sycophancy in large language models — Dubois, Ududec, Summerfield, Luettgau (UK AISI), Evaluating whether AI models would sabotage AI safety research — Kirk, Souly, Fronsdal, D'Cruz, Davies (UK AISI) (April 2026 UK AISI companion papers extending specific trends-report findings)
- instance-of: AI Benchmarks and Evaluation