This page summarizes two independent retrospectives that grade Leopold Aschenbrenner's June 2024 *Situational Awareness* essay against real-world data. Nathan Delisle's Situational Awareness: A One-Year Retrospective (LessWrong, June 2025) is a quantitative audit of the "drivers and indicators" behind the OOM-per-year scaling thesis. Jamie_Harris's How did Leopold do? Evaluating Situational Awareness's predictions (EA Forum, 2026-03-29) is an eight-category evaluation spanning infrastructure, capabilities, government response, China dynamics, safety, and the AGI-2027 timeline. Across both, the infrastructure and compute forecasts are graded on-track or ahead, benchmark capabilities on-track, revenue off by roughly 1.7x, the forecast nationalization not yet materializing, Chinese competitive dynamics and open-source under-modeled, the security critique vindicated, and the AGI-by-2027 timeline unresolved.
Sources
- Nathan Delisle, Situational Awareness: A One-Year Retrospective, LessWrong, June 2025
- Jamie_Harris, How did Leopold do? Evaluating Situational Awareness's predictions, EA Forum, 2026-03-29
Summary of findings
Both retrospectives converge on the same split. Aschenbrenner's physical forecast — compute, capex, power, chips, and algorithmic efficiency — is graded roughly right. His political forecast — nationalization, zero-sum US-China dynamics without independent Chinese innovation, and the collapse of open-source at the frontier — is graded partially wrong.
On infrastructure, the buildout is ahead of forecast on capex and on-track on chips and power, with the largest single training cluster slightly behind the projected waypoint. On capabilities, the benchmark trajectory is met but the qualitative "shocking leap" did not land. On political economy, the race framing is graded vindicated while the nationalization mechanism is graded wrong; independent Chinese innovation and the continued health of open-weight models at the frontier are the two structural misses. On safety and security, the shift from "alignment safety" toward "model/weight security" as the operative political concern tracks Aschenbrenner's framing. On multilateral coordination, the US and China opting out of the February 2026 responsible-military-AI declaration is graded consistent with Aschenbrenner's prediction that race dynamics would preempt coordination.
Scorecard
The combined scorecard, with Aschenbrenner's forecast, the actual state as of March 2026, each retrospective's verdict, and the grading source:
| Prediction | Forecast | Actual (March 2026) | Verdict | Source |
|---|---|---|---|---|
| Compute scaling ~0.5 OOM/yr | 0.5 OOM/yr | Mixed across labs; Grok 3 +0.38 OOM; OpenAI/Anthropic -0.2–0.5 OOM | On-pace in aggregate | Delisle |
| Algorithmic efficiency ~0.5 OOM/yr | 0.5 OOM/yr | 0.4–0.8 OOM/yr observed | Ahead | Delisle |
| Post-training ("unhobbling") effective compute | Major contributor | ~0.56 OOM/yr effective | Confirmed | Delisle |
| 2025 hyperscaler capex | ~$300B | ~$320B disclosed | On-pace | Delisle |
| Accelerator shipments 2025 | 12–16M units | In-range | On-pace | Delisle |
| Power demand | ~3% of US grid | 2–3 GW committed, ~3% of grid | On-pace | Delisle |
| Largest single training cluster | ~300 MW waypoint | Meta ~100k H100s, xAI Colossus — ~1/4 OOM behind | Slightly behind | Delisle |
| HBM supply | Not specifically flagged | Binding bottleneck | Missed (minor) | Delisle |
| AI revenue run-rate by mid-2026 | $100B | ~$60B | Behind (~1.7x low) | Harris |
| Benchmark capability trajectory | Aggressive | "Broadly met" | On-target | Harris |
| Qualitative "shocking leap" / wake-up moment | Expected | Did not land | Miss | Harris |
| Agentic "drop-in coworker" | By ~2025–27 | Tool use on-track; drop-in not realized | Partial | Harris |
| US government "Project" (nationalization) | By 2027–28 | No nationalization; export controls + AISI buildout instead | Behind / wrong mechanism | Harris |
| US-China strategic contest | Inevitable | Confirmed — espionage cases, $2.5B chip smuggling, 7nm CN fab, dueling framings | Confirmed on specifics | Harris |
| China = distillation only | Implied | Wrong — DeepSeek, Qwen, Kimi all produced genuine innovation | Miss | Harris |
| Open-source collapses at frontier | Implied | Wrong — open-weight models thrive at frontier | Miss | Harris |
| Lab security inadequate | High-confidence claim | Still inadequate; state-actor exploitation confirmed | Vindicated | Harris |
| "Safety → security" framing | Implicit | Industry rebrand tracks | Confirmed | Harris |
| Multilateral responsible-AI-military declarations | Will fail under race pressure | US + China opt out of Feb 2026 declaration | Confirmed | Harris |
| AGI by ~2027 | Central thesis | Unresolved; "more credible than six months ago" | Open | Harris |
Methodology
Delisle's audit is purely quantitative, tracking compute, capex, chips, power, and revenue waypoints. Harris graded the eight prediction categories using Claude for primary grading with Gemini red-teaming, a two-model evaluation pipeline. Neither author is Aschenbrenner; both are third-party evaluations rather than self-grading. This contrasts with the AI 2027 self-evaluation, which is first-party (Source: blog.aifutures.org).
Relation to other pages
The primary source under evaluation is Situational Awareness — The Decade Ahead; the retrospectives grade its broad trajectory as validated, with revenue, nationalization, and China/open-source structure as the notable misses. The assessment also feeds the track-record of Leopold Aschenbrenner. The one-year-later evaluation of AI 2027 (Source: blog.aifutures.org) follows a parallel retrospective structure on a parallel forecast.
On AGI Timelines, the AGI-by-2027 prediction remains unresolved but is graded "more credible than six months ago." On AI Race Dynamics, the race framing is graded vindicated, with the US and China opting out of the February 2026 military-AI declaration as recent evidence. For Compute Governance, the infrastructure trajectory supports the aggressive-scaling assumptions behind compute-threshold legislation. The Fast-Follow Problem is most directly affected by the structural miss on Chinese independent innovation (DeepSeek, Qwen, Kimi). The retrospectives also provide partial support for the skeptic side of AI as Normal Technology: no "shocking leap," the drop-in coworker not realized, and a government response slower than predicted.
Relationships
- supports: Situational Awareness — The Decade Ahead (partial — on infrastructure, capabilities, security, race framing)
- contradicts: Situational Awareness — The Decade Ahead (partial — on revenue, nationalization, China-distillation-only, open-source collapse)
- related: Grading AI 2027 (Source: blog.aifutures.org) — parallel retrospective structure on a parallel forecast
- related: Leopold Aschenbrenner
- related: AI Race Dynamics, Compute Governance, Fast-Follow Problem
- related: AGI Timelines