AI and productivity refers to the empirical evidence on how AI adoption affects worker productivity, task allocation, and output quality. The evidence comes from several structurally different vantage points — task-level field experiments, sentiment surveys, clinical endpoints, enterprise-portfolio surveys, provider-side observational data, and internal lab disclosures — that measure different quantities and do not always agree. Documented effects range from 14–26% within-deployment productivity gains in customer support and software development to near-zero measurable P&L impact across most enterprise pilots, alongside a separately documented erosion of unaided human performance (deskilling). It is a key input to the labor disruption timeline debate.
Snapshot
Quantitative findings, newest source on top. Each row cites its source; the analytical context appears in the thematic sections below.
| Date / source | Finding | Source |
|---|---|---|
| Jun 2026 — Anthropic "When AI Builds Itself" | 8× code merged per engineer per day in Q2 2026 vs. 2021–2024; >80% of merged production code Claude-authored (May 2026); ~4× self-reported output (median, Mar 2026 poll of 130 research-team staff); ~800 fixes in Apr 2026 cutting a class of API errors ~1,000× (est. 4 human-years) | When AI Builds Itself (The Anthropic Institute, June 4, 2026) |
| Jun 2026 — Demirer et al., NBER w35275 | Files created or edited tripled; software releases rose ~30%; no increase in downloads of new apps | (Source: nber.org) |
| Mar 2026 — Anthropic Economic Index | Users with 6+ months Claude tenure show 3–4 pp higher success rates than newer users (after controls); +7 pp more work use; +6% input education level; education required by prompts rises ~1 year per additional year of Claude use | Anthropic Economic Index — March 2026: Learning Curves |
| Jan/Feb 2026 — Anthropic Economic Index | Top-10 task share on Claude.ai fell 24% → 19%; US top-5 state share fell 30% → 24%; Gini coefficient declining (parity projection slowed from 2–5 yrs to 5–9 yrs); API vs Claude.ai directive 64% vs 32%, work-focused 74% vs 46%, per-task success 49% vs 67%; coding 35% of Claude.ai conversations | Anthropic Economic Index — January 2026: Economic Primitives, Anthropic Economic Index — March 2026: Learning Curves |
| Jan 2026 — Anthropic Economic Index (macro extrapolation) | Claude usage alone projected to contribute 1.2–1.8 pp annual productivity growth over the next decade (0.6–0.8 pp under complements; 2.2–2.6 pp under substitutes) | Anthropic Economic Index — January 2026: Economic Primitives |
| 2026 — AMA Physician AI Sentiment Report (N=1,692) | Physician AI use 38% (2023) → 81% (2026); use cases per physician 1.1 → 2.3; top tasks research summaries 39%, discharge/care-plan generation 30%, chart documentation 28%; 78% expect efficiency help (vs 7% harm); 70% see burnout-task offload; 73% administrative offload; burnout-helpful sentiment 48% (2024) → 64% (2026); 88% concerned about skill erosion (trainees 70%; early-career 35% vs late-career 24%); privacy 13% helpful vs 41% harmful | AMA Physician AI Sentiment Report (2026) |
| 2026 — Stanford HAI AI Index | 14–26% productivity gains in customer support and software development; AI agent deployment in single digits across nearly all business functions; developers aged 22–25 saw employment fall ~20% from 2024 | Stanford HAI AI Index Report 2026 |
| 2026 — Anthropic 81K User Survey (N=80,508) | Productivity most common positive experience (32.0%); professional excellence top vision category (18.8%); 18.9% reported AI hasn't delivered | What 81,000 People Want from AI |
| 2025 — MIT Project NANDA, State of AI in Business | ~95% of enterprise GenAI pilots show no measurable P&L impact after est. $30–40B spend through mid-2025; ~5% of custom tools reach production; vendor partnership ~67% vs internal build ~33% success; >50% of budgets target sales/marketing; >90% of employees use personal LLM accounts; 66% of executives want feedback-learning systems, 63% want context retention; method: 300+ initiatives, 52 executive interviews, 153 survey responses | MIT NANDA — The GenAI Divide (State of AI in Business 2025) |
| 2025 — Lancet colonoscopy study | Non-AI adenoma detection rate fell 28.4% → 22.4% after endoscopists began regular AI use | Lancet Endoscopist Deskilling Study (2025) |
| 2023 — Brynjolfsson, Li, Raymond | 15% average productivity increase (issues resolved/hour) across 5,172 customer-support agents | Generative AI at Work |
Background and scope
Studies of AI's productivity effects measure several distinct quantities that are often conflated. Task-level field experiments measure within-deployment gains at one job, firm, or tool. Sentiment surveys measure worker perception and self-reported efficiency. Clinical-endpoint studies measure unaided-baseline performance over time. Enterprise-portfolio surveys measure how often pilots reach production at all. Provider-side observational data report what users of a deployed system actually do. Internal-lab disclosures report productivity in the most AI-saturated knowledge-work environments. Because these measure different things, headline numbers that appear to conflict — large per-task gains versus near-zero enterprise P&L impact — are frequently compatible once the quantity each measures is specified.
Task-level field evidence
Two field studies anchor the case that AI raises within-deployment productivity, both finding the largest gains among lower-skilled or lower-ability workers.
In customer support, Brynjolfsson, Li, and Raymond (2023) found a 15% average increase in issues resolved per hour across 5,172 agents (Generative AI at Work). Gains were largest for the least-skilled workers, with AI acting as a "leveler"; the most experienced workers saw small gains in speed but small declines in quality. AI also facilitated worker learning, with new agents improving faster when given AI access, and customer experience improved, with customers more polite and fewer escalations.
In software development, Hoffmann et al. (2025) found that GitHub Copilot shifted developers toward core coding and away from project management (Generative AI and the Nature of Work). It increased autonomous rather than collaborative work and exploration rather than exploitation, with effects greater for lower-ability individuals — which the authors read as AI flattening organizational hierarchies.
OpenAI reported a usage-side finding in its first Work at the Frontier study, published July 31, 2026: 43.5% of occupation-specific ChatGPT messages involve tasks traditionally associated with a different occupation. The figure is the company's own and describes message content rather than measured output, so it bears on task boundaries rather than on productivity directly (Source: openaiglobalaffairs.substack.com).
The Stanford HAI AI Index (2026) reports 14–26% productivity gains in customer support and software development, with weaker or negative effects in tasks requiring more judgment (Stanford HAI AI Index Report 2026). The same source notes that AI agent deployment remains in single digits across nearly all business functions, and that developers aged 22–25 saw employment fall roughly 20% from 2024 even as older-developer headcount grew.
Enterprise-portfolio evidence
Where the task-level studies measure gains within deployments that work, MIT Project NANDA's State of AI in Business 2025 measures how often enterprise GenAI pilots reach that stage at all (MIT NANDA — The GenAI Divide (State of AI in Business 2025)). It finds that roughly 95% of enterprise GenAI pilots show no measurable P&L impact after an estimated $30–40 billion in enterprise spend through mid-2025, with only about 5% of custom enterprise AI tools reaching production. The study reviewed 300+ public initiatives, conducted 52 executive interviews, and collected 153 senior-leader survey responses.
NANDA's other findings cut against common assumptions. Vendor partnerships succeeded roughly twice as often as internal builds (about 67% versus 33% success rates), inverting the "build for differentiation" intuition. Budgets and returns were misaligned: more than 50% of GenAI budgets targeted sales and marketing, while back-office automation yielded the highest returns. More than 90% of employees reported using personal LLM accounts for work ("shadow AI"), so employee adoption outpaced formal enterprise deployment. NANDA identifies the core barrier as a "learning gap": most deployed tools are stateless and do not retain feedback, adapt to context, or improve over time. 66% of executives wanted systems that learn from feedback and 63% wanted context retention.
A practitioner account published in July 2026 reaches a stronger conclusion than NANDA's from a much smaller and non-random base. Consultant Nikhil Suresh reports that every AI project his firm has observed over eighteen months has failed, and argues that "almost every report at a company about 'massive AI productivity gains' is untrue as a matter of brute fact," while allowing that genuine gains exist as the exception (AI Mania Is Eviscerating Global Decision-Making (Nikhil Suresh)). He identifies two reporting distortions that would inflate self-reported figures independently of any real effect: engineers who complete work without AI reporting that they used it, because managers are dissatisfied otherwise; and staff measured on token consumption running models in self-prompting loops and discarding the output. In one client engagement his team found that staff were unaware they had been issued AI licences at all, which he treats as undermining the productivity claims attached to them. The account is first-person, its sources deliberately anonymised, and it should be read as attributed observation rather than measurement.
NANDA and the task-level studies are not contradictory because they measure different quantities. NANDA measures the denominator — what share of enterprise AI attempts becomes a successful deployment — while Brynjolfsson, Copilot, and the AI Index measure the numerator's per-deployment payoff. AI raises productivity substantially when deployed well; successful enterprise deployment is rare. The full mechanism — learning gap, integration friction, workflow-redesign burden, vendor-versus-build, shadow AI — is treated in Enterprise AI Deployment Gap.
Healthcare sector evidence
The AMA Physician AI Sentiment Report (2026) provides sector-wide deployment-reality data from a high-skill clinical profession outside software (AMA Physician AI Sentiment Report (2026)). Physician AI use roughly doubled in three years, from 38% in 2023 to 81% in 2026 (N=1,692), and average use cases per physician rose from 1.1 to 2.3. As in other professions, gains concentrate in documentation-heavy tasks: research summaries (39%), discharge and care-plan generation (30%), and chart documentation (28%) — the clinical analogue of customer-support ticket resolution and developer boilerplate.
Perceived efficiency and burnout effects were large and positive: 78% expected AI to help work efficiency (versus 7% expecting harm), 70% saw AI as offloading burnout-contributing clinical tasks, and 73% saw administrative offload. Burnout-helpful sentiment rose from 48% to 64% between 2024 and 2026. A counter-signal ran alongside: 88% of physicians were at least mildly concerned about AI-induced skill erosion, with concern highest for trainees (70%) and concentrated among early-career physicians (35% versus 24% for late-career) — the group whose skills are still being built through practice. That subjective concern aligns with the clinical-endpoint data in Lancet Endoscopist Deskilling Study (2025), where the non-AI colonoscopy adenoma detection rate fell from 28.4% to 22.4% after endoscopists began regular AI use. Privacy was the only dimension where physicians expected net harm (13% helpful versus 41% harmful), a regulatory-salience signal distinct from the efficiency narrative.
Provider-side observational evidence
The Anthropic Economic Index adds provider-side observational data from a frontier lab reporting on its own deployed traffic (Anthropic Economic Index — January 2026: Economic Primitives, Anthropic Economic Index — March 2026: Learning Curves). Anthropic sampled 1M Claude.ai conversations plus 1M first-party API records at two points (Nov 13–20, 2025 and Feb 5–12, 2026), classified each along five "economic primitives" (task complexity, skills, use case, autonomy, success), and released the data — the first source in which the lab itself, using its own classifier, reports what its users do.
Over the short horizon (Nov 2025 → Feb 2026), the top-10 task share on Claude.ai fell from 24% to 19%, indicating diversification of workloads rather than concentration, and the US top-5 state share fell from 30% to 24% of per-capita usage, with the Gini coefficient continuing to decline (though the parity projection slowed from 2–5 years to 5–9 years). Claude.ai and the API behaved as distinct populations: the API was more concentrated, more directive (64% versus 32%), more work-focused (74% versus 46%), and lower per-task success (49% versus 67%) — consistent with production automation versus human-in-the-loop use. Coding was 35% of Claude.ai conversations, the single dominant category, with evidence that coding workloads are migrating from Claude.ai to API-based agent tooling (see Agentic AI).
The March 2026 report added the first published first-party user-tenure analysis. After controls for task type (O*NET category and request cluster), model selection, country, language, and use case, users with 6+ months of tenure on Claude showed 3–4 pp higher success rates than newer users. They also used Claude for more work (+7 pp), with more complex inputs (+6% education level), and in more collaborative rather than directive modes, and the education required by prompts rose roughly one year per additional year of Claude use. This is a different productivity channel than Brynjolfsson's "leveler" finding: Brynjolfsson is a within-day comparison showing AI lifts low-skill workers toward high-skill levels, while the Anthropic tenure result is a within-user, over-months slope showing skilled usage compounds. Both can hold simultaneously, jointly predicting that AI initially levels (newest users benefit most on simple tasks) then stratifies (users who invest tenure pull ahead on complex tasks) — consistent with skill-biased technological change rather than purely equalizing dynamics, and with Acemoglu's capital-labor concerns discussed below.
Anthropic's January report estimates that Claude usage alone would contribute 1.2–1.8 pp of annual productivity growth over the next decade (0.6–0.8 pp under complements assumptions; 2.2–2.6 pp under substitutes). These are Anthropic's own model-based extrapolations from observed speedups and task-time estimates, not independently measured productivity gains, and they are not directly comparable to Acemoglu's ≤0.66%/10-year TFP ceiling, which is whole-economy.
Several limitations qualify how the Index should be read. Success is classified by Claude evaluating whether Claude succeeded, and the January report acknowledges classifier-version variance (Sonnet 4.5 versus Sonnet 4); NANDA's 95% pilot-failure finding measures P&L impact rather than conversation-level task completion, a harder bar. Anthropic also chooses what counts as a task, what counts as success, and which composition metrics to report — composition data (coding share, geographic distribution, Claude.ai versus API split) is hard to game and is new information, whereas success rates are provider-graded and the productivity extrapolations are Anthropic's scenario. The Index cannot observe Lancet-style deskilling by design, since it measures with-AI performance only; the Anthropic learning curve (with-AI performance improving over time) and the Lancet finding (without-AI performance degrading over time) are fully compatible observations of the same workers. Finally, emerging API workflows such as sales enablement and automated trading doubled in three months, which might appear to contradict NANDA's pilot-failure finding but does not: Anthropic measures traffic across all paths, including vendor-mediated ones, while NANDA measures organizations attempting internal builds, and NANDA itself finds vendor-partnership paths succeed roughly twice as often. As a citing rule of thumb, the composition data is the most reliable, the success-rate data should be read as provider-graded, and the productivity extrapolations as Anthropic's scenario rather than measurement.
Internal-lab evidence and the release gap
A narrower but higher-end view is Anthropic's own AI-building-AI productivity, disclosed in "When AI Builds Itself" (4 June 2026). Anthropic reports 8× code merged per engineer per day in Q2 2026 versus 2021–2024, with more than 80% of merged production code Claude-authored as of May 2026, and roughly 4× self-reported output (median, March 2026 poll of 130 research-team staff). It also cites work that would not otherwise have happened, including about 800 fixes in April 2026 that cut a class of API errors roughly 1,000×, an estimated four human-years of work.
The essay self-caveats more than the Economic Index: lines-of-code "almost certainly overstates the true productivity gain," developer self-estimates are known to overshoot (it cites METR's own work on this), and Amdahl's law means that speeding code generation relocates the bottleneck — human code review is now Anthropic's binding constraint. The data, from the most AI-saturated knowledge-work environment that exists, resolves to faster output with a moving bottleneck rather than unbounded acceleration.
An MIT-led NBER working paper (Demirer et al., w35275, June 2026) provides a downstream counterpoint: AI coding gains are not flowing through to shipped software. Files created or edited tripled, but software releases rose only about 30%, with no increase in downloads of new apps (Source: nber.org). The two findings are compatible — Anthropic measures upstream code-production throughput inside one elite lab, while the NBER paper measures downstream release and adoption across the broader developer economy, where review, integration, QA, and distribution remain human-paced. Together they illustrate a recurring pattern in which per-task AI gains are large but their conversion into realized output is gated by the slowest un-automated stage. See Recursive Self-Improvement (RSI) for the upstream loop and Enterprise AI Deployment Gap for the downstream gate.
User perspectives
The Anthropic 81K User Survey (What 81,000 People Want from AI) is the largest direct user account of AI productivity, covering 80,508 Claude users. Productivity was the most commonly reported positive experience (32.0%), described as AI that "dramatically sped up work and automated repetitive tasks." The top vision category was professional excellence (18.8%): wanting AI to handle mundane tasks so users can focus on strategic, higher-level work. At the same time, 18.9% reported that AI had not delivered on their expectations, one respondent noting that "AI should be cleaning windows and emptying the dishwasher so I can paint and write poetry. Right now it's exactly the other way around." Many goals labeled "productivity" masked deeper quality-of-life desires, such as leaving work on time, cooking dinner, or spending time with family. The sample is self-selected (Claude.ai users who accepted the interview) and will overrepresent positive experiences, but it offers qualitative texture on what "productivity" means to users that task-level field studies abstract away.
Augmentation, displacement, and deskilling
The evidence points in more than one direction. On augmentation, productivity gains are largest for least-skilled workers (both task-level studies), workers shift toward higher-value activities (the Copilot study), and AI facilitates learning and improves the work experience. On displacement, entry-level employment is declining in the same fields that show productivity gains (HAI Index); software engineer Matt Shumer wrote in a personal account that he was "no longer needed for the actual technical work of my job" (Source: shumer.dev); and Dario Amodei predicts that half of entry-level white-collar jobs could be disrupted in one to five years. One reading is that augmentation comes first and displacement later, as AI initially makes workers more productive and improving capabilities then let firms achieve the same output with fewer workers.
A third channel, introduced by the Lancet 2025 colonoscopy study, is captured by neither the augmentation nor the displacement frame: AI deskilling, the erosion of unaided human performance while the worker remains employed and the combined human-plus-AI team performs normally. Deskilling is invisible to standard productivity metrics because they measure combined output; only the counterfactual — the AI turned off, unavailable, or contraindicated — reveals it. This has three implications for the productivity literature. Task-level studies may overstate durable human capability: Brynjolfsson's 15% and Stanford HAI's 14–26% measure combined output, so if the unaided baseline is drifting downward, human-capital stock is eroding beneath the headline gain. Field-experiment comparators may not be stable: trials using "unaided clinician" or "unaided developer" as a control assume that baseline is time-invariant, but chronic AI exposure may make it time-varying in the direction that flatters the intervention arm. And resilience costs are mis-measured: effective productivity in a world with occasional AI downtime, regulatory removal, or contraindicated edge cases is lower than task-level gains imply.
The macro skeptic view
Economist Paul Kedrosky (May 2026) offers the strongest current empirical counterpoint to the productivity-gains literature, arguing that AI is "nowhere to be seen yet in any meaningful productivity data anywhere" and appears only in non-residential fixed investment at levels comparable to the railroad build-out or rural electrification — large upfront investment with long and uncertain diffusion timelines (Source: AI Is Really Weird). Kedrosky's claim differs from Acemoglu's model-based 0.66% ceiling in being empirical: measured productivity statistics do not yet show AI's impact. It is consistent with NANDA's 95% enterprise pilot-failure finding and with the Narayanan/Kapoor AI as Normal Technology view that diffusion takes decades, and it is in tension with task-level studies (Brynjolfsson, HAI Index) and Anthropic's deployment data, which measure within-deployment gains rather than economy-wide statistics.
Steve Newman reached a similar conclusion in a July 20, 2026 essay, "Anecdotes Everywhere, Evidence Almost Nowhere", arguing that AI has so far had little clearly measurable macroeconomic impact. He cited a St. Louis Fed estimate that AI-related investment added 0.97 percentage points to US GDP growth in the first three quarters of 2025 — an investment effect rather than a productivity effect — while noting that the figure sums all spending on information-processing equipment, software, R&D, and data-center construction, so it works as an AI estimate only if non-AI spending in those categories was flat year over year. He treated the Census Bureau finding that about 18% of firms had adopted AI by year-end 2025 as uninformative, since it records only that a firm checked a box saying it uses AI "in any of [their] business functions." On labor he cited the Stanford "Canaries in the Coal Mine" finding of 16% lower employment for early-career workers in the most AI-exposed occupations, bounding it as covering about 7% of workers studied, measured relative to similar-age workers in less-exposed jobs, with absolute employment down under 10% and no broad employment effect visible even among software engineers. His summary position is that "absence of evidence is suggestive evidence of absence" for the period through roughly mid-2026, paired with the argument that 2026 may be the last year for which that holds, given AI metrics growing at 3× to 10× per year. See AI Labor Disruption.
Reconciling micro and macro
Brynjolfsson et al.'s +15% task-level gain and Acemoglu's ≤0.66% TFP-over-10-years macro ceiling are often cited as opposing views but are not contradictory; both can be true simultaneously because they measure different things.
| [[generative-ai-at-work | Brynjolfsson, Li, Raymond (2023)]] | [[simple-macroeconomics-of-ai | Acemoglu (2024)]] | |
|---|---|---|---|---|
| What's measured | Task-level productivity at one job, one firm, one AI tool | Aggregate TFP across the whole economy | ||
| Headline | +15% issues/hour | ≤0.66% TFP over 10 years | ||
| Method | Empirical rollout across 5,172 agents | Task-based macro model applying Hulten's theorem | ||
| Scope | Immediate, single-task | 10-year economy-wide projection |
A 15% gain in a single task translates into aggregate GDP growth only after weighting by that task's share of total output, netting out complementarities, and accounting for the fact that early evidence comes from easy-to-learn tasks while future gains must come from hard-to-learn ones. Acemoglu's methodological argument is that bullish forecasts (Goldman, McKinsey) extrapolate task-level results mechanically to the whole economy, which is not how GDP moves.
Both papers are consistent with the micro finding that AI gains accrue most to lower-skilled workers; what they do not resolve is the distributional question at the macro level. On the equalizing view, Brynjolfsson's "leveler" finding, if it generalizes, implies AI could compress labor-income inequality by lifting low-skill productivity toward high-skill levels. On the capital-labor view, Acemoglu predicts AI will widen the gap between capital and labor income regardless of its within-labor distribution, because task-level savings accrue to firms and compute rather than workers. Both can hold at once — labor-income inequality narrowing while the capital-labor gap widens — which is the contested empirical question rather than whether productivity is up.
Two further macro sources bear on this. Autor (2024) frames AI as a potential tool to rebuild middle-skill expertise, conditional on how it is deployed; this aligns with the "leveler" micro findings but does not resolve the capital-labor question (Applying AI to Rebuild Middle-Class Jobs). The 2028 Global Intelligence Crisis scenario introduces the concept of "Ghost GDP" — output that appears in national accounts but never circulates through the real economy because machines do not spend on discretionary goods. If AI productivity gains flow primarily to capital and compute rather than labor, the macro effects could be deflationary even as measured productivity rises (Ghost GDP framing).
Federal Reserve Board staff took up the same reconciliation problem in a July 2026 FEDS note (The AI Buildout and the Economy: Publicly Available Data to Assess AI's Impact (FEDS Notes, July 2026)) that proposes a framework of public indicators for tracking the buildout across capabilities and costs, investment and adoption, and productivity and labor. Tracking sectoral labor productivity by AI exposure, it finds that "productivity trends across all three levels have been relatively consistent over time, suggestive of micro-level productivity gains not adding up in aggregate," and offers four candidate explanations rather than one: micro experiments measure task-level rather than firm-level output, so "a 10% improvement on a task does not necessarily lead to proportional gains for a firm if adjustment costs or other bottlenecks lie elsewhere"; gains observed in information-sector roles may not generalize; historically "measured productivity gains from GPTs lag investment by years"; and measurement itself may misattribute gains between capital deepening and total factor productivity, with services output inferred from revenue distorted by price declines.
The note also separates technical feasibility from economic deployment in a way the benchmark literature generally does not. METR's task-completion horizon "measure[s] technical feasibility, but not cost-effective deployment," which turns on a fixed cost of adjustment — integration "into firm-specific systems," which the authors describe as "not directly observable and difficult to quantify" — and the marginal cost of inference, which is. Its overall assessment as of 2026 is that the evidence "is consistent with a buildout phase rather than the onset of broad-based displacement" (The AI Buildout and the Economy: Publicly Available Data to Assess AI's Impact (FEDS Notes, July 2026)).
Related concepts
- AI Labor Disruption — the broader risk framework
- AI Deskilling — erosion of unaided performance; distinct from displacement or augmentation
- Enterprise AI Deployment Gap — the divergence between task-level capability and enterprise-portfolio success
- Labor Disruption Timelines — where these findings fit in the timeline debate
- AI as Normal Technology — the augmentation evidence supports the slow-diffusion thesis
- Agentic AI — the shift from augmentation to autonomous task completion