AI Policy Wiki
Dashboard

Grading AI 2027's 2025 Predictions (AI Futures Project, February 2026)

high confidence · updated 2026-07-25

Self-assessment by Eli Lifland and Daniel Kokotajlo of the AI 2027 scenario's 2025 predictions. Reports aggregate quantitative progress at roughly 65% of the predicted pace (58–66% by category), most qualitative predictions on pace, and an implied takeoff window shifted to mid-2028 through mid-2030 after adjusting for expected compute and labor growth slowdowns.

"Grading AI 2027's 2025 Predictions" is a February 12, 2026 post by Eli Lifland and Daniel Kokotajlo on the AI Futures Project blog, assessing how the first year of the *AI 2027* scenario compares to what happened. It is a self-assessment by the scenario's own authors, and reports both where the scenario ran ahead of events and where the authors judge their original presentation to have been in error.

The headline finding: "In aggregate, progress on quantitative metrics is at roughly 65% of the pace that AI 2027 predicted." Most qualitative predictions are assessed as on pace.

Method

The authors describe the exercise as a fourth forecasting method complementing those already in use, and state it in four steps: build a detailed, concrete trajectory; wait; check whether events are roughly on track or veering off entirely; and if roughly on track, quantitatively estimate the pace of real progress against the scenario, then adjust the forecast correspondingly.

For each quantitative prediction they estimate a "pace of progress" multiplier where 1× is on pace, 2× twice as fast, and 0.5× half as fast. Aggregating by prediction category gives 58–66%; aggregating over individual predictions gives a higher figure (mean 75%, median 84%), which they explicitly reject as the worse indicator because 7 of the 15 individual predictions concern compute, over-weighting compute forecasts relative to capability indicators — and, they note, most of their timeline uncertainty concerns what capability level a given amount of compute buys, which capability indicators track directly.

Implied takeoff window

AI 2027 depicted a takeoff from full coding automation to superintelligence over the course of 2027. At 65% of the depicted rate, that takeoff would run from late 2027 to mid-2029. The authors then apply a further adjustment for expected slowdowns in training-compute and human-labor growth — a lower effective-compute growth rate, before accounting for AI R&D automation — using the AI Futures Model, which moves the window to mid-2028 through mid-2030. They note the model does not account for hardware R&D automation, which would shorten its takeoff estimates.

The authors record that this window does not match either of their personal medians. Kokotajlo's median for full coding automation is 2029 — later than mid-2028 — while the implied two-year takeoff to superintelligence is slower than his roughly one-year median. Lifland's median for full coding automation is in the early 2030s, with a takeoff median of about two years. They also note that over the course of 2025 their timelines lengthened. See AGI Timelines.

Quantitative findings

MetricAssessment
SWE-Bench VerifiedBehind. AI 2027 predicted 85% by mid-2025 from a 72% base; the best actual score was 74.5% (Opus 4.1). The 2025 Epoch forecasting survey showed the same directional error, with a median forecast of 88% by end-2025 against an actual 81%.
Coding time horizon (METR 80%)Ambiguous, and the ambiguity is the authors' own. Progress is at 1.04× a central AI-2027-speed trajectory from the April 2025 timelines model, but 0.66× the trajectory actually displayed on the published graph, which contained an error. The authors state that predictions made with their newer model would fall between 0.66 and 1.04.
RevenueSlightly ahead. OpenAI's annualized revenue reached about $20B against an $18B prediction. Forecasters in the Epoch survey underpredicted the sum of AGI-company revenues by roughly 2×.
ValuationWell behind. OpenAI was at $500B as of October 2025, up from $300B at AI 2027's publication; the scenario had $500B valuations arriving in June 2025.
AI software R&D upliftBehind, primarily because the authors revised their estimate of early-2025 uplift downward, leaving end-2025 estimates close to the scenario's original start-of-2027 figures.
Compute growthMostly on pace, with the possible exception of the largest training run. The authors estimate that no leading company has conducted a substantially larger run than GPT-4.5 (February 2025), while stating "extremely wide uncertainty" given the obscurity around training compute.

Qualitative findings

The authors grade the scenario's prose predictions passage by passage.

Mid-2025 agents. The depiction of computer-using "personal assistant" agents that check in for confirmation and struggle to reach widespread usage is assessed as correct, with ChatGPT agent (July 2025) and its Expedia booking demonstration cited as the analogue of the scenario's DoorDash example. The prediction that specialized coding and research agents would begin transforming their professions is judged "fairly accurate," anchored on Anthropic's September 2025 statement that Claude Code had passed $500 million in run-rate revenue with usage growing more than 10× in three months. The authors qualify one detail: they do not think there was an especially large amount of agent usage in Slack or Teams, though they hold the spirit of the prediction correct. On reliability they note it is "possible that coding agents were slightly more reliable than we expected."

The competitive field. The scenario's fictional OpenBrain was imagined with competitors 3–9 months behind. The authors report the race is closer: "more like a 0-2 month lead between the top US AGI companies."

AI-assisted AI research. Assessed as partially correct — AI is helping substantially with coding but less with other parts of AI research, though the authors note the scenario did not claim AIs would be great at all of AI research.

Continuous retraining. The prediction that "finishes training" would become a misnomer is judged correct, citing the outside inference that GPT-4o, GPT-5, and GPT-5.1 are probably different continuations of the same base model — noting this is guessed by outsiders rather than confirmed — and the generally more frequent release pace.

Dangerous capabilities. Hacking assistance to humans is assessed as "very strong," with the authors noting it is unclear how capable AIs are on their own. Bioweapon capabilities are judged on track, citing OpenAI's upgrade of its bio capability level to High and Anthropic's upgrade to ASL-3. See CBRN Uplift and AI and Cybersecurity.

Model specs and alignment. The scenario's description of a written model specification combining vague goals with specific dos and don'ts, memorized and reasoned over through AI-assisted training toward helpfulness, harmlessness, and honesty, is graded correct but discounted: "This was already true at the time we published. It remains true now, but as predictions go, this was an easy one." See OpenAI Model Spec.

Extreme incidents. AI 2027 predicted that real deployments would produce no incidents as extreme as Gemini telling a user to die or Bing Sydney's 2023 behavior. The authors volunteer a potential counterexample against themselves — Grok's July 2025 "MechaHitler" episode — and note that their original footnote 27 restricted the prediction to incidents not deliberately prompted by a user, leaving it unclear how much the episode should count, since it combined user-prompted and autonomous behavior.

Metrics the authors commit to tracking

Four indicators are named for 2026, each with the scenario benchmark attached:

  1. AI R&D uplift studies and surveys. AI 2027 depicted 1.9× software R&D uplift by end-2026. The authors cite METR's randomized controlled trial of early-2025 AI coding tools on experienced open-source developers, whose headline result was a slowdown, against Anthropic's survey of its own technical staff returning a median 2× coding uplift — which, they note, still implies well under 2× uplift for AI software R&D as a whole because of compute bottlenecks.
  2. AGI company revenues and valuations. The scenario depicted the leading company at $55B annualized revenue and a $2.5T valuation by 2026.
  3. Coding time horizon. A central AI-2027-speed trajectory predicts roughly three-work-week 80% time horizons by end-2026; the newer AI Futures Model's handcrafted AI-2027-speed trajectory reaches about a year. The authors note time horizons will get harder to measure as models improve.
  4. Other benchmarks. They note that, apart from coding time horizon, AI 2027 registered no predictions for benchmarks that did not yet exist when it was written, and express hope that higher-difficulty benchmarks arrive in 2026.

Their closing assessment is that even with these indicators, "it will unfortunately be difficult to be highly confident by the end of 2026 that AI takeoff will or won't begin in 2027."

Relationships