AI Policy Wiki
Dashboard

When AI Builds Itself (The Anthropic Institute, June 4, 2026)

high confidence · updated 2026-06-06

Anthropic Institute essay arguing AI is already accelerating AI development — >80% of merged production code is Claude-authored, 8× code per engineer vs. 2021–2024 — laying out three futures (stall-and-diffuse, compounding efficiency, full recursive self-improvement) and arguing the world should build the verification infrastructure needed for the option of a coordinated frontier slowdown or pause.

Primary text: anthropic.com Authors: Marina Favaro and Jack Clark, The Anthropic Institute (editorial support: Santi Ruiz). Published June 4, 2026.

"When AI Builds Itself" is an essay published June 4, 2026 by Marina Favaro and Jack Clark of The Anthropic Institute, with editorial support from Santi Ruiz. It argues that AI is already meaningfully accelerating AI research and development, drawing on public benchmarks together with previously-unreported internal Anthropic data; situates that trend on the path toward recursive self-improvement (RSI); sketches three forward scenarios; and concludes with a policy commitment that the Anthropic Institute will work to build the verification systems a credible coordinated slowdown or pause would require. The authors frame it as an argument for building the infrastructure for a verifiable coordinated slowdown or pause of frontier AI development rather than for pausing now.

Core argument

The essay's central claim is conditional: "We believe it would be good for the world to have the option to slow or temporarily pause frontier AI development." It does not call for an immediate pause. The authors argue that a unilateral pause "accomplishes much less" because it changes who leads rather than the underlying dynamics, and that an effective pause would require multiple well-resourced labs in multiple countries stopping under the same conditions, each able to verify that the others have actually stopped. Building that verification capability, not pausing now, is the action item the authors identify.

Evidence that AI is accelerating AI development

From public benchmarks, the authors argue the rate of capability gain is itself increasing. The length of tasks models can reliably complete on their own is doubling roughly every four months, up from an earlier roughly seven-month trend. The essay illustrates this with software tasks: Claude Opus 3 (March 2024) handled roughly 4-minute tasks; Claude Sonnet 3.7, about a year later, reached roughly 90-minute tasks; and Claude Opus 4.6 reached roughly 12-hour tasks. Extrapolated, the authors expect days-long tasks "this year" and weeks-long tasks in 2027. On SWE-bench (real bug-fixing on real open-source codebases), scores went from low single digits to saturation in about two years. On CORE-Bench (reproducing the results of a published paper, a prerequisite for original research), scores went from about 20% in 2024 to saturation about 15 months later. On the METR long-task benchmark, Claude Mythos Preview worked for "at least" 16 hours and was "at the upper end of what [METR] can measure without new tasks."

The essay's novel contribution is internal Anthropic data. As of May 2026, more than 80% of code merged into Anthropic's codebase was authored by Claude, up from low single digits before Claude Code's February 2025 research-preview launch. A footnote distinguishes this conservative lines-merged-to-production figure from leadership's public "90%+ of code" claim, which includes scripts and experimental code. Code merged per engineer per day in Q2 2026 was 8× the 2021–2024 level; the figure was flat for the first four years, climbed in 2025 when Claude began running code rather than just suggesting it, and steepened again in 2026 with longer-horizon autonomy. The authors caveat that lines-of-code overstates the true productivity gain. A March 2026 poll of 130 Anthropic research-team employees put median self-reported output at roughly 4× what it would have been without AI, which the authors call likely an overestimate but directionally corroborated.

The essay cites specific work it argues would not otherwise have happened: in April 2026 Claude shipped 800+ fixes that cut a class of API errors by roughly 1,000×, which the overseeing engineer estimated at roughly 4 years of human work. On code quality, Anthropic staff judged Claude-written code worse than human code in late 2025, "roughly at parity today," and expected it to be "strictly better within the year"; an automated Claude code-reviewer would have caught roughly one-third of the bugs behind past claude.ai production incidents. On experiment optimization, on a fixed "make this training code run faster" task, Claude went from roughly 3× speedup (Opus 4, May 2025) to roughly 52× (Mythos Preview, April 2026), against a skilled human who needs 4–8 hours to reach roughly 4×. On open-ended research, in April 2026 Claude-powered agents recovered 97% of a weak-to-strong-supervision gap over roughly 800 cumulative agent-hours and roughly $18,000 of compute, versus roughly 23% for two human researchers over roughly a week, with caveats that the result did not transfer cleanly to production-scale models and that the humans chose the problem and scoring rubric. On research judgment, across n=129 real Claude Code sessions, the best November 2025 model (Opus 4.5) beat the human's next-step choice 51% of the time, and Mythos Preview (April 2026) reached 64%, on moments deliberately chosen where the human's move had room to improve.

The authors' throughline is that Claude is at or beyond human level at execution (engineering, running specified experiments) but still lags at direction-setting and research taste (choosing which problems and experiments matter). That gap, they write, is "the gap between AI today and a future system that could autonomously design its own successor."

Three scenarios

The essay sketches three forward scenarios. In the first, the trend stalls but today's capabilities diffuse widely: exponentials turn out to be S-curves, whether because taste does not come from scaling, because the binding constraint is compute, energy, or supply chain, or because of an exogenous shock. The authors include this scenario "for completeness" but consider it unlikely, noting no measured capability curve has bent yet; it gives society the most time to adapt. They note diffusion continues regardless, citing Project Glasswing finding 10,000+ high or critical vulnerabilities.

In the second scenario, compounding efficiency gains occur but humans still set direction: AI development becomes substantially automated while humans keep choosing problems and judging results, and 100-person firms do the work of 10,000–100,000. The authors say this would revolutionize knowledge work and government services but would also enable authoritarian surveillance and individualized influence operations at scale, and they consider it the likeliest path. They invoke Amdahl's law, under which speeding one stage just moves the bottleneck; they note human code review is already Anthropic's new bottleneck.

In the third scenario, full recursive self-improvement, AI builds its successors: pace becomes governed by compute and by algorithmic-efficiency discovery, and humans move to oversight, validation, and verification of an expanding "virtual lab." The authors say the alignment outcome is what they are "least certain about" — models could be aligned enough to find novel solutions or wise enough to halt, or rare present-day misalignment could compound until control is lost. They again invoke Amdahl's law to bound real-world change, writing that more intelligence "can't learn what a drug does over decades of use, can't hold elections sooner than a constitution dictates."

Verification and pause-infrastructure argument

The essay's distinctive policy content concerns why an effective pause is hard, in the authors' framing. It would need multiple frontier labs in multiple countries stopping under the same triggering conditions, each able to verify the others. Detectability is harder for AI than for prior arms-control regimes, the authors argue: training runs are far easier to conceal than missile silos, inputs are general-purpose, and the incentive to defect quietly is enormous because whoever continues inherits the lead. A credible pause, they write, "has to specify what triggers it, what lifts it, and who adjudicates." The authors point to precedent in the Intermediate-Range Nuclear Forces Treaty and other verification regimes, but note those "took decades to build both the infrastructure and the trust. We don't have that long."

The commitment is that the Anthropic Institute will research and help build the verification systems a credible slowdown or pause would require, and will convene policymakers, researchers, civil society, and other labs "in the coming months." The argument rests on Compute Governance, the main enforcement and detectability theory, and sits at the most-binding end of the instrument spectrum described in Frontier AI Governance. The essay is a lab-internal, conditional, and verification-first position, distinct from prohibition proposals such as ControlAI's ASI Security Bill and the 30+ Canadian parliamentarians' "trust but verify" call, which are legislative rather than lab-internal.

The essay supplies internal-lab figures bearing on Recursive Self-Improvement (RSI) — the more-than-80% code figure, the 8× figure, the roughly 3× to roughly 52× experiment-optimization range, and the 97% gap recovery — alongside benchmarks. Earlier coverage of recursive self-improvement drew on OpenAI's GPT-5.3 Codex acknowledgment and on Dario Amodei's qualitative "1–2 years" remark in The Shape of the Thing.

Reception

The essay was covered by The Wall Street Journal, Reuters, and SiliconANGLE on June 4, 2026. Gary Marcus published a "no need to panic" response the same day. It appeared three days after Anthropic's confidential IPO filing, and critics including David Sacks have previously framed Anthropic's safety posture as a "regulatory capture agenda." This reception is documented further on Anthropic and is not part of the primary text.

Relationships

Sources

Primary text retrieved verbatim from anthropic.com (the Anthropic Institute essay). Pulled and authenticity-verified by the 2026-06-05 gap scan against the canonical anthropic.com host; corroborated by WSJ, Reuters, and SiliconANGLE (all 2026-06-04). Verification record: Wiki/queue/gap-scan/proposed-sources/anthropic-when-ai-builds-itself-2026.md.