AI Policy Wiki
Dashboard

Remote Labor Index: Measuring AI Automation of Remote Work (Mazeika et al., 2025)

high confidence · updated 2026-07-26

Introduces the Remote Labor Index, a multi-sector benchmark of real-world, economically valuable projects measuring end-to-end agent performance rather than research-oriented capability. Finds AI agents perform near the floor, with the highest-performing agent reaching a 2.5% automation rate.

"Remote Labor Index: Measuring AI Automation of Remote Work" is a benchmark paper from the Center for AI Safety and Scale AI, with a large author list including Mantas Mazeika, Alice Gatti, Summer Yue, Alexandr Wang, Bing Liu, and Dan Hendrycks.

The measurement gap

The paper's premise is a disconnect between benchmark progress and economic effect: "AIs have made rapid progress on research-oriented benchmarks of knowledge and reasoning, but it remains unclear how these gains translate into economic value and automation."

It names the resulting problem for policy: "we lack standardized, empirical methods for monitoring the trajectory of AI automation." Without such methods, claims about displacement rest on capability benchmarks that were never designed to measure economic substitution.

The benchmark

The Remote Labor Index (RLI) is "a broadly multi-sector benchmark comprising real-world, economically valuable projects designed to evaluate end-to-end agent performance in practical settings."

Two design choices distinguish it from capability benchmarks. The tasks are real projects with economic value rather than constructed problems, and evaluation is end-to-end — the agent must deliver a complete piece of work rather than demonstrate a component skill.

The result

"AI agents perform near the floor on RLI, with the highest-performing agent achieving an automation rate of 2.5%."

The gap between this figure and performance on research benchmarks is the paper's substantive finding: the same systems that score highly on knowledge and reasoning evaluations complete very few real projects end to end. The stated purpose is to "ground discussions of AI automation in empirical evidence, setting a common basis for tracking AI impacts."

Bearing on the labor debate

The 2.5% figure is a measurement of current end-to-end automation, not a forecast, and the benchmark is designed to be tracked over time — its value is as a series rather than a point. It supplies a direct empirical answer to the question Newman treats as unresolved, and sits alongside the Stanford early-career finding as evidence of the gap between demonstrated capability and realized displacement. It bears on the near-term phase of the staged frameworks in Anthropic's Economic Policy Framework (June 2026) and A Roadmap for the Upcoming Labor Transition (Cheng & Schaal, June 2026), both of which assume displacement is not yet broadly visible.

Relationships