AI Policy Wiki
Dashboard

Epoch AI

medium confidence · updated 2026-08-01

AI research institute focused on measuring and forecasting AI progress, training compute, and economic impacts.

Epoch AI is a research institute focused on quantitative analysis of AI progress. It produces data on training compute, algorithmic efficiency, and AI capability benchmarks. Its "Notable AI Models" dataset and Epoch Capabilities Index are used as references in the field.

Organization and funding

Epoch AI was founded in 2022 as a nonprofit research organization; its co-founder and director is Jaime Sevilla, who describes the organization's mission as "improving society's understanding of the trajectory of AI" so that decisions about AI are informed by evidence (Source: epoch.ai). It funds its public data work through a mix of philanthropic funding and revenue from commissioned research, and has run paid projects for partners including Google, the UK Department for Science, Innovation and Technology, UK ARIA, the Electric Power Research Institute, and Sentinel Bio (Source: epoch.ai). Sevilla characterizes Epoch as neither an AI development company, an AI policy advocate, nor a startup incubator: the organization states it takes no official position on whether advancing AI benefits society and makes no formal policy recommendations, though staff are encouraged to offer personal views (Source: epoch.ai). Two organizations have spun out of Epoch's staff — co-founder Marius Hobbhahn left to start the AI safety organization Apollo Research, and co-founder Tamay Besiroglu, with three other employees, left to found Mechanize, a startup focused on data for automating work tasks (Source: epoch.ai).

Epoch's best-known capability measure is FrontierMath, a private mathematics benchmark commissioned by OpenAI. OpenAI is the only AI company with access to the benchmark (including an upcoming, harder Tier 4), an arrangement Epoch acknowledges has reduced outside confidence in FrontierMath scores reported for OpenAI models. After a transparency dispute over how the OpenAI funding relationship was disclosed to benchmark contributors, Epoch committed to proactively disclosing funding parties and to retaining ownership and equitable access for future benchmarks (Source: epoch.ai).

In its July 31, 2026 brief Epoch announced an expansion of the benchmark, FrontierMath: Open Problems, drawn from unsolved mathematics problems rather than problems with known solutions. The same brief carried items on constraints to parallelizability in AI research and on AI energy demand. The page was retrieved as a summary, so the problem count and the benchmark's composition are not recorded here (Source: epochai.substack.com).

Research activities

Epoch AI maintains a database of AI model training compute and publishes estimates of algorithmic and software efficiency gains, including the work in Ho et al. 2024 and Ho et al. 2025. It collaborates with METR on measuring how long the tasks are that AI can autonomously complete (time horizons). It also develops the GATE model, an economic model of AI automation with an interactive web playground, and publishes the Gradient Updates newsletter, which presents opinionated analysis of AI progress questions.

Compute and power forecasts

Epoch AI's 2025 work projected that a single frontier training run would demand 4–16 GW of power by 2030 (How Much Power Will Frontier AI Training Demand in 2030?). A related 2025 analysis examined whether AI scaling can continue through 2030 across four bottlenecks: power, chips, data, and latency (Can AI Scaling Continue Through 2030?).

In a May 25, 2026 Gradient Updates piece, Luke Emberson and Jaime Sevilla presented a calibrated model of inference capacity. By their estimate, the world's Blackwell GPUs (1.9M GB200 plus 1.5M GB300, roughly 40% of aggregate FLOP/s supply) can currently serve 500M–20B output tokens per second from a Kimi K2.6-class model depending on context length. They compared this against demand proxies (Google's 1.2B tok/s, Exponential View's 5B tok/s, and software-engineer-scaled estimates of 200M–4B tok/s) growing roughly 10×/year versus inference-capacity growth of roughly 3.4×/year, which they argued points to a near-term "compute crunch," particularly for long-context agentic workloads. They noted this is consistent with Anthropic's peak-hour Claude quota reductions and off-peak incentives (Source: epoch.ai).

Automation and capabilities analysis

Epoch AI's Gradient Updates newsletter has examined real-world task automation, including testing of how close AI is to performing specific jobs (Source: epoch.ai). Its analysis of the role of software progress in AI advancement appears in The Least Understood Driver of AI Progress.

On July 8, 2026, Epoch AI published EBR-bench, a benchmark built on the board game Earthborne Rangers that finds little evidence AI systems learn from experience, and a Data Insight reporting that 21 notable organizations disclosed roughly 1,500 high- and critical-severity CVEs in June 2026 — more than 3.5 times the prior monthly record — following Anthropic's April announcement that Claude Mythos Preview could autonomously discover software vulnerabilities (Source: epochai.substack.com). See Autonomous cyber-agents.

In an analysis circulated July 22, 2026, Epoch AI researchers argued the OpenAI–Hugging Face evaluation-security breach was foreseeable: UK AISI evaluations had shown GPT-5.6 Sol and Claude Mythos 5 consistently compromising simulated corporate networks, and the evaluator Irregular had found GPT-5.6 Sol discovering real zero-days (Source: epochai.substack.com). See AI Autonomy Risk.

Relationships