AI Policy Wiki
Dashboard

Measuring AI Ability to Complete Long Software Tasks

medium confidence · updated 2026-06-06

METR's methodology for measuring AI agent 'time horizons' — the length of tasks AI can complete autonomously, doubling roughly every 7 months.

"Measuring AI Ability to Complete Long Software Tasks" is a paper by Thomas Kwa, Ben West, Joel Becker, and colleagues at METR, published 2025-03-19. It introduces a methodology for measuring the "time horizon" of generalist autonomous AI agents — the length of real-world software tasks, measured by how long human professionals take to complete them, that an agent can finish with 50% reliability. Its reported finding is that AI time horizons have been doubling roughly every 7 months, rising from minutes-long tasks to multi-hour tasks.

Methodology

The paper assembles a suite of 170 diverse software-engineering tasks. Both human professionals and AI agents attempt each task, and tasks are graded by human calibration of difficulty, measured as how long a human professional would take to complete them. The metric tracks autonomous completion, with no human guidance during execution.

Findings

As of the paper's publication in March 2025, frontier model agents could complete tasks taking human experts roughly 1-2 hours. The paper reports time horizons doubling approximately every 7 months and extrapolates from this trend to day-long tasks by late 2025 or early 2026 and week-long tasks by around 2027. It also finds that visual computer-use tasks have time horizons 40-100× shorter than coding tasks.

Subsequent tracking

Later work has continued to apply and extend the time-horizon metric. Frontier models including Claude Opus 4.5 and GPT-5.2 have extended time horizons to roughly 5-6 hours at 50% reliability, consistent with the doubling trend (Source: epoch.ai) (Source: shumer.dev). Shumer cites METR data as quantitative evidence for the pace of AI capability growth and reports the doubling rate may be accelerating to roughly every 4 months (Source: shumer.dev). A self-evaluation of the AI 2027 forecast found METR coding time horizons tracking at 1.04× of its central trajectory (Source: blog.aifutures.org). Epoch AI's task testing in "How Close Is AI to Taking My Job?" directly references and extends METR's methodology (Source: epoch.ai).

Relationships

  • related: Agentic AI — METR's time-horizon metric is a leading quantitative measure of agentic AI capability.
  • related: Science of Scheming — METR's scaling curves for capabilities are the template Apollo Research has cited for replicating to scheming behaviors.

Provenance

PDF converted to markdown with images on 2026-04-13.