AI Policy Wiki
Dashboard

Situational Awareness: One-Year-Later Retrospectives

medium confidence · updated 2026-06-06

Two independent retrospectives (Delisle June 2025 quantitative audit; Harris March 2026 eight-category evaluation) grade Leopold Aschenbrenner's 2024 Situational Awareness thesis — infrastructure and compute roughly on-track, revenue ~1.7x low, 'Project' nationalization not materializing, China/open-source dynamics underweighted, security concerns vindicated.

This page summarizes two independent retrospectives that grade Leopold Aschenbrenner's June 2024 *Situational Awareness* essay against real-world data. Nathan Delisle's Situational Awareness: A One-Year Retrospective (LessWrong, June 2025) is a quantitative audit of the "drivers and indicators" behind the OOM-per-year scaling thesis. Jamie_Harris's How did Leopold do? Evaluating Situational Awareness's predictions (EA Forum, 2026-03-29) is an eight-category evaluation spanning infrastructure, capabilities, government response, China dynamics, safety, and the AGI-2027 timeline. Across both, the infrastructure and compute forecasts are graded on-track or ahead, benchmark capabilities on-track, revenue off by roughly 1.7x, the forecast nationalization not yet materializing, Chinese competitive dynamics and open-source under-modeled, the security critique vindicated, and the AGI-by-2027 timeline unresolved.

Sources

  1. Nathan Delisle, Situational Awareness: A One-Year Retrospective, LessWrong, June 2025
  2. Jamie_Harris, How did Leopold do? Evaluating Situational Awareness's predictions, EA Forum, 2026-03-29

Summary of findings

Both retrospectives converge on the same split. Aschenbrenner's physical forecast — compute, capex, power, chips, and algorithmic efficiency — is graded roughly right. His political forecast — nationalization, zero-sum US-China dynamics without independent Chinese innovation, and the collapse of open-source at the frontier — is graded partially wrong.

On infrastructure, the buildout is ahead of forecast on capex and on-track on chips and power, with the largest single training cluster slightly behind the projected waypoint. On capabilities, the benchmark trajectory is met but the qualitative "shocking leap" did not land. On political economy, the race framing is graded vindicated while the nationalization mechanism is graded wrong; independent Chinese innovation and the continued health of open-weight models at the frontier are the two structural misses. On safety and security, the shift from "alignment safety" toward "model/weight security" as the operative political concern tracks Aschenbrenner's framing. On multilateral coordination, the US and China opting out of the February 2026 responsible-military-AI declaration is graded consistent with Aschenbrenner's prediction that race dynamics would preempt coordination.

Scorecard

The combined scorecard, with Aschenbrenner's forecast, the actual state as of March 2026, each retrospective's verdict, and the grading source:

PredictionForecastActual (March 2026)VerdictSource
Compute scaling ~0.5 OOM/yr0.5 OOM/yrMixed across labs; Grok 3 +0.38 OOM; OpenAI/Anthropic -0.2–0.5 OOMOn-pace in aggregateDelisle
Algorithmic efficiency ~0.5 OOM/yr0.5 OOM/yr0.4–0.8 OOM/yr observedAheadDelisle
Post-training ("unhobbling") effective computeMajor contributor~0.56 OOM/yr effectiveConfirmedDelisle
2025 hyperscaler capex~$300B~$320B disclosedOn-paceDelisle
Accelerator shipments 202512–16M unitsIn-rangeOn-paceDelisle
Power demand~3% of US grid2–3 GW committed, ~3% of gridOn-paceDelisle
Largest single training cluster~300 MW waypointMeta ~100k H100s, xAI Colossus — ~1/4 OOM behindSlightly behindDelisle
HBM supplyNot specifically flaggedBinding bottleneckMissed (minor)Delisle
AI revenue run-rate by mid-2026$100B~$60BBehind (~1.7x low)Harris
Benchmark capability trajectoryAggressive"Broadly met"On-targetHarris
Qualitative "shocking leap" / wake-up momentExpectedDid not landMissHarris
Agentic "drop-in coworker"By ~2025–27Tool use on-track; drop-in not realizedPartialHarris
US government "Project" (nationalization)By 2027–28No nationalization; export controls + AISI buildout insteadBehind / wrong mechanismHarris
US-China strategic contestInevitableConfirmed — espionage cases, $2.5B chip smuggling, 7nm CN fab, dueling framingsConfirmed on specificsHarris
China = distillation onlyImpliedWrong — DeepSeek, Qwen, Kimi all produced genuine innovationMissHarris
Open-source collapses at frontierImpliedWrong — open-weight models thrive at frontierMissHarris
Lab security inadequateHigh-confidence claimStill inadequate; state-actor exploitation confirmedVindicatedHarris
"Safety → security" framingImplicitIndustry rebrand tracksConfirmedHarris
Multilateral responsible-AI-military declarationsWill fail under race pressureUS + China opt out of Feb 2026 declarationConfirmedHarris
AGI by ~2027Central thesisUnresolved; "more credible than six months ago"OpenHarris

Methodology

Delisle's audit is purely quantitative, tracking compute, capex, chips, power, and revenue waypoints. Harris graded the eight prediction categories using Claude for primary grading with Gemini red-teaming, a two-model evaluation pipeline. Neither author is Aschenbrenner; both are third-party evaluations rather than self-grading. This contrasts with the AI 2027 self-evaluation, which is first-party (Source: blog.aifutures.org).

Relation to other pages

The primary source under evaluation is Situational Awareness — The Decade Ahead; the retrospectives grade its broad trajectory as validated, with revenue, nationalization, and China/open-source structure as the notable misses. The assessment also feeds the track-record of Leopold Aschenbrenner. The one-year-later evaluation of AI 2027 (Source: blog.aifutures.org) follows a parallel retrospective structure on a parallel forecast.

On AGI Timelines, the AGI-by-2027 prediction remains unresolved but is graded "more credible than six months ago." On AI Race Dynamics, the race framing is graded vindicated, with the US and China opting out of the February 2026 military-AI declaration as recent evidence. For Compute Governance, the infrastructure trajectory supports the aggressive-scaling assumptions behind compute-threshold legislation. The Fast-Follow Problem is most directly affected by the structural miss on Chinese independent innovation (DeepSeek, Qwen, Kimi). The retrospectives also provide partial support for the skeptic side of AI as Normal Technology: no "shocking leap," the drop-in coworker not realized, and a government response slower than predicted.

Relationships