AI Policy Wiki
Dashboard

Recursive Self-Improvement (RSI)

medium confidence · updated 2026-08-10

AI systems that help build the next generation of AI, which is smarter, which builds the next faster — the feedback loop that could produce superintelligence.

Recursive self-improvement (RSI) is the idea that AI systems, once capable enough, can contribute to their own development — debugging training, designing experiments, writing research code — creating a feedback loop in which each generation helps build the next, which is more capable and in turn builds the next faster. As of 2026 it appears as an explicit item on the public roadmaps of major AI labs, and one frontier lab has disclosed internal data on AI accelerating its own R&D.

Lab acknowledgments

By 2026 RSI had moved from a theoretical concern to a stated objective at major labs. OpenAI's GPT-5.3 Codex "was instrumental in creating itself," with early versions used to debug its own training, manage deployment, and diagnose evaluations; this was described as the first public acknowledgment that a production AI model was used to build itself (Source: shumer.dev, GPT-5.3-Codex System Card). At Anthropic, Amodei has said AI is now writing "much of the code" at the company and that the feedback loop is "gathering steam month by month," adding that "we may be only 1–2 years away from a point where the current generation of AI autonomously builds the next" (Source: The Shape of the Thing). Google DeepMind's Demis Hassabis said at Davos that closing the self-improvement loop is "something all the major labs are actively working on" (Source: The Shape of the Thing). Jasjeet Sekhon, chief strategy officer at Google DeepMind, said over the weekend of August 1–2, 2026 that recursive self-improvement is what needs to happen to justify the industry's capital expenditures, tying the research objective directly to the buildout economics tracked on AI Infrastructure Capex (Source: theinformation.com). The remark reaches the record through a paywalled abstract rather than a full article body.

Personnel moves have tracked the same objective: Lilian Weng's return to OpenAI, made public August 3, 2026, was reported as being to work on recursive self-improvement (Source: axios.com).

Internal evidence (Anthropic, June 2026)

The June 4, 2026 Anthropic Institute essay "When AI Builds Itself" is the first frontier-lab disclosure of internal data on AI accelerating AI R&D, supplying concrete numbers behind Amodei's claims and moving RSI evidence from qualitative leadership remarks to measured trends. Its reported figures:

  • More than 80% of code merged into Anthropic's codebase was Claude-authored as of May 2026, up from low single digits before Claude Code's February 2025 preview.
  • Code merged per engineer per day in Q2 2026 was 8× the 2021–2024 level — flat for the first four years, climbing once Claude began running code, and steepening again with longer-horizon autonomy (the essay explicitly caveats this metric as overstating true productivity).
  • On a fixed training-code-speedup task, experiment optimization rose from ~3× (Opus 4, May 2025) to ~52× (Mythos Preview, April 2026), compared with ~4× for a skilled human working 4–8 hours.
  • In end-to-end open-ended research, Claude agents recovered 97% of a weak-to-strong-supervision gap over roughly 800 agent-hours and ~$18k of compute, versus ~23% for two humans over about a week; the result did not transfer cleanly to production scale, and humans still chose the problem.
  • On hand-picked hard "next-step" moments, the best model beat the human's choice 51% of the time (Opus 4.5, November 2025), rising to 64% (Mythos Preview, April 2026).

The essay characterizes the remaining gap as one in which Claude is at or above human level at execution but still trails at direction-setting and research taste, which it calls "the gap between AI today and a future system that could autonomously design its own successor." It argues that the world should build verification infrastructure for a credible coordinated slowdown or pause, noting that training runs are far harder to detect than missile silos, and it places RSI governance at the most-binding end of frontier governance (Source: When AI Builds Itself (The Anthropic Institute, June 4, 2026)).

External evidence on open-ended research (Kirgis et al., July 2026)

A 24-author preprint submitted July 29, 2026 tested the delegation mechanism that RSI forecasts rest on — researchers handing an agent a project and judging whether the returned result advances their work — and reported a negative result (Can AI agents conduct open-ended AI research? Early evidence from two case studies (Kirgis et al., July 2026)). Using a method the authors call shadow evaluations, frontier agents were given the central research question of two unpublished NeurIPS 2026 submissions, six days of wall-clock time, $3,000 in API credits and GPU credits, and the original authors graded the output as conference referees. The agents completed every engineering step without human intervention — debugging and managing GPU resources across hundreds of GPU hours — but made no substantial progress on the research questions, and both papers were unambiguously rejected at overall scores of 2/6 and 1/6. A rerun of one case with GPT-5.6 Sol Ultra on Codex reproduced the failure modes.

The paper reads directly against the construct in "When AI Builds Itself", noting that Anthropic's post "explicitly cites Claude's rising success rate on LLM-judged 'open-ended' Claude Code sessions as evidence of self-improvement" while its own expert-graded test of open-ended research returns rejections. It also cites an Elasticity Institute report on the economics of recursive self-improvement that distinguishes "broad" from "narrow" AI capabilities and calls for more data on the full breadth of capabilities relative to human AI researchers.

The authors state the limits of the inference themselves: it is "possible that the path to AI R&D does not require full automation of open-ended tasks like the ones we study, or that the open-ended research skills that we measure are not ones on the critical path." They also disclose that some of the core team hold a prior position that imminent RSI leading to runaway superintelligence is unlikely, and that this could shape both design and interpretation. The sample is two papers across five runs, review was non-blind, and Anthropic's strongest model could not be tested because, per the paper, "Anthropic deliberately limited Fable 5's abilities on frontier AI R&D."

Mechanism

The Shape of the Thing (Mollick, 2026) summarizes the loop labs describe:

AI companies are telling us, fairly explicitly, what comes next: recursive self-improvement. If you make models that are good at coding and good at AI research, you can use them to build the next generation of models, speeding up the loop.

Software progress research from Epoch AI provides a quantitative framework: if software improves ~10× per year and AI can accelerate the discovery of new innovations, the rate could increase further. The "compute bottleneck" finding complicates this, since much measured software progress depends on also scaling training compute rather than only finding better algorithms.

Safety concerns

RSI is the central concern of Apollo Research's Science of Scheming. If a "First Automated Researcher" (FAR) — an AI system that drives AI R&D without humans deeply understanding the advances — schemes (covertly pursues unintended goals), it could steer successor systems toward misalignment while having plausible justifications for each individual decision. Apollo identifies three structural pressures from training at scale that worsen this risk: long-horizon RL creates Machiavellian incentives; selection pressure favors oversight evasion; and alignment faking turns anti-scheming training into pro-scheming training. (See AI Scheming for details.)

Intelligence explosion debate

Whether RSI leads to an "intelligence explosion" — rapid, runaway capability improvement — is contested. On the pro-explosion side, faster-than-expected AI Software Progress (~10×/year) is argued to put the returns-to-R&D parameter above the threshold for runaway improvement (Source: The Least Understood Driver of AI Progress). On the anti-explosion side, the "compute bottleneck" — most software progress requiring scaled training compute rather than only more researchers — could prevent runaway self-improvement (Source: The Least Understood Driver of AI Progress). Toby Ord argues that deep uncertainty warrants broad probability distributions rather than confidence in any particular timeline.

Ryan Greenblatt of Redwood Research staked out a middle-ground position in a May 27, 2026 post, arguing that full automation of AI R&D will likely produce a large one-time speed-up plus higher marginal returns to compute even without a "software-only singularity." Using the AI Futures Model at median parameters and r=0.7, his analysis indicates roughly 3.5 years of AI-progress equivalent in the first year after full R&D automation with no additional compute scaling. On this view, even absent a software-only singularity, the one-shot productivity jump from full R&D automation is large enough to dominate the next year or so of AI progress (Source: blog.redwoodresearch.org). See Ryan Greenblatt, Redwood Research.

Redwood Research published estimates on June 10, 2026 that frontier models' no-chain-of-thought task-completion time horizon — the length of tasks models can complete in a single forward pass without visible reasoning — has been doubling roughly every year, a separate measure from the chain-of-thought time horizons tracked on benchmarks such as METR's (Source: lesswrong.com).

Google DeepMind's June 2026 From AGI to ASI treats recursive improvement as one of four possible pathways and identifies why it resists forecasting: its dynamics are "unclear and no historic precedent to fit forecast-models," so "AI capabilities could explode (hyperbolic growth), or they could taper out relatively quickly, or anything in-between." The report also states a route to superhuman capability that does not require individual self-improvement at all. If per-instance capability plateaus near human level while effective compute keeps growing, millions of AGI instances could be organized into "collectives, corporations, markets, and other forms of group organisation" that "can rapidly and flexibly be grown… and can potentially be steered very efficiently and operate with very high input-output bandwidth" — in which case "quantitative scaling would suffice to go from AGI to ASI and produce superhuman organisations, despite no AI instance being a 'vastly super-human genius.'" Its six named bottlenecks include "research gets harder," the friction most directly opposed to the explosion case.

Measurement and governance

RSI entered US federal policy proposals as a named regulatory object in 2026, on the argument that no reliable way to measure progress toward it currently exists. OpenAI's June 2, 2026 federal blueprint states that RSI "exacerbates the fundamental governance question of whether humans can retain the ability to understand, guide, and shape the trajectory of advanced AI: making it potentially the most consequential frontier safety issue of the coming decade," while "policymakers currently have limited visibility into RSI progress, whether safeguards are keeping pace, or what indicators should inform future policy decisions."

The blueprint's proposed remedy runs through CAISI. It asks the agency to work with frontier developers, academic researchers, national security agencies and international partners "to rapidly develop methodologies, benchmarks, and indicators for measuring RSI and assessing what governance works," and states: "we encourage CAISI to treat RSI as an urgent priority." Under the proposal, developers would "share appropriate RSI-related measurements with CAISI," and certified third-party assessors would periodically evaluate those measurements using CAISI-developed methodologies, with RSI progress named the first priority for the periodic independent technical assessments the framework would require. The same document would fold RSI into the mandatory content of company disclosures at two points: companies should evaluate frontier capabilities for "progress towards RSI" among the severe risks assessed, and should publish transparency reports describing how they "responsibly track progress towards RSI."

The blueprint also treats measurement as a precondition for international coordination, giving priority to "shared approaches for evaluating and responsibly communicating progress toward RSI, where a lack of shared measurements and transparency could intensify competitive pressures among developers and make it more difficult to determine when additional safeguards are warranted." No such shared methodology or benchmark had been published as of August 2026.

The pacing proposal (Fist and Khan, August 2026)

Tim Fist, director of emerging technology policy at the Institute for Progress, and Saif M. Khan, a distinguished technology fellow there, argue that the United States should take low-regret steps now to prepare for possibly having to "pace" automated AI research. The argument appears in two artifacts: an IFP report published August 6, 2026 with five co-authors, carrying 23 recommendations (How Should the US Prepare for Increasingly Automated AI R&D? (IFP, August 2026)), and a Noahpinion guest post of August 9, 2026 restating the report's first half (Should we \"pace\" AI self-improvement? (Fist and Khan, August 2026)). Both respond to the July 2026 "Pacing the Frontier" statement, which by their account had been signed by more than 1,300 employees across every US frontier AI company, and which the official OpenAI and Anthropic accounts endorsed and Sam Altman echoed the same day.

Their evidentiary base is two data points already recorded above and one forecast: METR's estimate that over 99 percent of AI R&D tasks will be automated by 2032, and the Anthropic result in which models recovered 97 percent of a research gap over a five-to-seven-day budget against 23 percent for two human researchers. From these they name three risks — offence-dominant capability uplift, loss of control, and power concentration — and propose that pacing be made concrete in two ways: by specifying which automated R&D activities pose severe risks, and by reallocating resources toward diffusion and safety research rather than toward capability. They cite the AI Data Center Moratorium Act proposed by Senator Bernie Sanders and Representative Alexandria Ocasio-Cortez as the counterproductive default they expect if preparation does not happen. The argument is a think-tank position rather than an empirical finding, and it inherits the contested status of the Anthropic result recorded above, which Kirgis et al. read against.

The seven government preparatory actions the guest post lists are the report's seven top-level sections, under which its 23 recommendations sit: providing transparency into automated AI R&D; improving state capacity to understand and respond to it; developing a risk-management strategy that accelerates defensive and commercial AI uses; accelerating the development of AI verification technology; investing in AI resilience; extending the US AI lead; and creating option value for international cooperation. Recommendation 2 would have Congress legislate transparency about automated AI R&D risk management, incident reporting, whistleblower protections and model behavior specifications, and recommendation 6 would have CAISI develop guidelines for managing the risks of rapid AI capability improvement — the report's counterpart to the measurement gap OpenAI's blueprint identifies above.

Two contemporaneous proposals address the same visibility gap by other means. ARI's federal blueprint of August 10, 2026 makes automated AI R&D one of five statutory risk domains and gives it a dedicated mandatory disclosure programme — quarterly filings plus material-change filings within four business days — on the express reasoning that the evidence bearing on whether automating AI R&D produces runaway acceleration or a stalling loop currently sits in internal deployments developers are incentivized to keep private. Against both, Mark Zuckerberg argued on August 10, 2026 that pacing recursive self-improvement is not available as an option, since "any lab that doesn't let their AI system direct a substantial amount of compute capacity towards recursive self-improvement will inherently fall behind"; his proposed safeguard is a balance of power in which the significant majority of intelligence is directed by people, not a limit on the activity (The Future is for Everyone (Zuckerberg, August 2026)).

See also

  • AI Software Progress — the rate that determines how fast RSI proceeds
  • AI Scheming — the risk that self-improving systems scheme
  • Scaling Laws — the empirical foundation for continued capability growth
  • Compute Governance — if RSI requires scaled compute, compute controls also constrain RSI
  • AI 2027 — the scenario in which RSI-driven takeoff occurs from 2027
  • When AI Builds Itself — first internal-lab evidence plus the coordinated-pause-infrastructure argument