A randomized controlled trial published by METR on 2025-07-12 (arXiv 2507.09089), authored by Joel Becker, Nate Rush, Elizabeth Barnes, and David Rein. The study measured how access to early-2025 AI coding tools affected the task-completion time of experienced developers working on mature open-source projects, and found that AI access increased completion time by 19 percent, contrary to the expectations of the developers themselves and of forecasters.
Findings
The study reported that AI assistance increased task-completion time by 19 percent for experienced open-source developers working in mature projects. This result ran counter to four separate expectations: developers' own pre-study predictions of being 24 percent faster, their post-study self-assessments of being 20 percent faster, economist predictions of 39 percent faster, and ML-specialist predictions of 38 percent faster. The gap between the measured slowdown and the post-hoc self-assessment indicates that participants believed AI had sped them up even though their measured times rose.
Methodology
The study used a randomized controlled trial (RCT) covering 16 experienced developers and 246 tasks drawn from mature open-source projects. Tasks were randomly assigned to an AI-permitted or an AI-restricted condition. Developers in the AI-permitted condition used Cursor Pro with Claude 3.5/3.7 Sonnet. The study measured actual completion times rather than relying on self-reported productivity.
The authors examined 20 potential explanatory factors, including project scale, quality standards, developer AI familiarity, and baseline speed. The 19 percent slowdown held across these analyses.
Limitations
The authors noted several limits to generalization: the sample is small (16 developers); participants were experienced developers, so AI may help novices more; the tasks came from mature projects, so AI may help greenfield work more; and the tooling reflected early-2025 capabilities, which may have changed since.
Reception and positioning
The study was the first rigorous RCT to report a result that contradicts the consensus that AI makes developers faster, and it has been used as a counter-citation to vendor productivity marketing and to internal-use claims from Anthropic and OpenAI, including the assertion that "100% of code written by AI." Its design has also been described as a methodological template for similar studies in other domains. The divergence between measured times and developer self-assessment has been characterized as a gap in developers' ability to judge whether AI is helping them on a given task.
Relationships
- supports: MIT NANDA — The GenAI Divide (State of AI in Business 2025), AI Is Really Weird, Why I Think AI Take-Off Is Relatively Slow (Cowen), Enterprise AI Deployment Gap.
- contradicts: Anthropic/OpenAI internal productivity claims; Generative AI at Work (which studied customer support, not software engineering); vendor benchmarks.
- depends-on: AI Benchmarks and Evaluation (jagged frontier framing).
- related: Measuring AI Ability to Complete Long Software Tasks (METR's other canonical work — time-horizon methodology).