"Remote Labor Index: Measuring AI Automation of Remote Work" is a benchmark paper from the Center for AI Safety and Scale AI, with a large author list including Mantas Mazeika, Alice Gatti, Summer Yue, Alexandr Wang, Bing Liu, and Dan Hendrycks.
The measurement gap
The paper's premise is a disconnect between benchmark progress and economic effect: "AIs have made rapid progress on research-oriented benchmarks of knowledge and reasoning, but it remains unclear how these gains translate into economic value and automation."
It names the resulting problem for policy: "we lack standardized, empirical methods for monitoring the trajectory of AI automation." Without such methods, claims about displacement rest on capability benchmarks that were never designed to measure economic substitution.
The benchmark
The Remote Labor Index (RLI) is "a broadly multi-sector benchmark comprising real-world, economically valuable projects designed to evaluate end-to-end agent performance in practical settings."
Two design choices distinguish it from capability benchmarks. The tasks are real projects with economic value rather than constructed problems, and evaluation is end-to-end — the agent must deliver a complete piece of work rather than demonstrate a component skill.
The result
"AI agents perform near the floor on RLI, with the highest-performing agent achieving an automation rate of 2.5%."
The gap between this figure and performance on research benchmarks is the paper's substantive finding: the same systems that score highly on knowledge and reasoning evaluations complete very few real projects end to end. The stated purpose is to "ground discussions of AI automation in empirical evidence, setting a common basis for tracking AI impacts."
Bearing on the labor debate
The 2.5% figure is a measurement of current end-to-end automation, not a forecast, and the benchmark is designed to be tracked over time — its value is as a series rather than a point. It supplies a direct empirical answer to the question Newman treats as unresolved, and sits alongside the Stanford early-career finding as evidence of the gap between demonstrated capability and realized displacement. It bears on the near-term phase of the staged frameworks in Anthropic's Economic Policy Framework (June 2026) and A Roadmap for the Upcoming Labor Transition (Cheng & Schaal, June 2026), both of which assume displacement is not yet broadly visible.
Relationships
- supports: AI Labor Disruption — supplies a direct end-to-end automation measurement against capability-benchmark inference
- supports: Anecdotes Everywhere, Evidence Almost Nowhere (Steve Newman, July 2026) — independent evidence for the claim that measurable displacement remains small
- related: AI Benchmarks and Evaluation — argues research benchmarks do not measure economic substitution
- related: AI and Productivity, Agentic AI, Dan Hendrycks, Scale AI