AI economic primitives are a five-dimension measurement framework introduced in the Anthropic Economic Index (AEI) January 2026 report for classifying AI interactions along axes intended to connect to traditional economic analysis of labor and productivity. The framework converts raw AI usage logs into economically interpretable signals, supporting trend analysis, cross-sector comparison, and connection to standard labor economics through O*NET occupational codes and BLS wage data.
The five primitives
Each AI conversation is classified along five dimensions.
| Primitive | What it measures | Example values |
|---|---|---|
| Task complexity | Estimated human hours required to complete the task | Software dev: 3.3 hrs avg; Personal tasks: 1.8 hrs avg |
| Skills required | Years of education implied by the task | Software dev: 13.8 yrs; Personal: 9.1–9.4 yrs |
| Use case | Work / Personal / Coursework classification | 46% Work (Claude.ai); 74% Work (API) |
| Autonomy | 1–5 scale (1 = highly supervised; 5 = fully autonomous) | ~3.5 avg across categories |
| Success rate | Proportion of tasks completed successfully | 61% (software dev); 78% (personal tasks) |
The framework was implemented using Claude itself as the classifier: a CLIO privacy-preserving pipeline analyzes conversations and assigns codes. This produces a self-referential measurement structure, a lab using its own model to measure its own usage, with advantages in scale and consistency and a limitation that model-graded success rates may not match user-perceived outcomes.
Empirical findings
Snapshot composition (November 2025, Jan 2026 report)
Drawn from 1M Claude.ai conversations and 1M API records (Anthropic Economic Index — January 2026: Economic Primitives):
- Coding and math tasks accounted for 34% of Claude.ai use and 46% of API use.
- Work-related use accounted for 46% of Claude.ai use and 74% of API use.
- Task success rates were 67% overall on Claude.ai and 49% overall on the API.
- Personal tasks succeeded more often than software-development tasks (78% vs. 61%). Anthropic identifies complexity, rather than domain, as the first-order success predictor.
The report frames the higher personal-task success rate as a counter to capability-centric narratives: in deployment, models are not straightforwardly better at coding than at personal tasks, and harder, more complex tasks fail more often regardless of domain.
Three-month trend (November 2025 to February 2026, Mar 2026 report)
The March 2026 report (Anthropic Economic Index — March 2026: Learning Curves) compared usage across two data points.
| Metric | Nov 2025 | Feb 2026 | Interpretation |
|---|---|---|---|
| Top-10 task share (Claude.ai) | 24% | 19% | Claude.ai diversifying |
| Top-10 task share (API) | 28% | 33% | API concentrating |
| Coding share (Claude.ai) | — | 35% | Dominant single category |
| Coursework share (Claude.ai) | 19% | 12% | Post-semester decline |
| Personal use share (Claude.ai) | 35% | 42% | Consumerizing |
| US top-5 states share | 30% | 24% | Geographic deconcentration |
| Avg. hourly wage of tasks | $49.30 | $47.90 | Slight decline |
Anthropic describes divergent patterns by surface: Claude.ai is diversifying and shifting toward personal use, broader task types, and slightly lower-skill tasks, while the API is concentrating, with fewer task types accounting for more traffic. Anthropic attributes the Claude.ai skill-mix decline partly to coding workloads migrating to API-based agent tooling such as Claude Code.
AI learning curves (Mar 2026 report)
The March 2026 report presents a user-tenure learning-curve analysis that Anthropic describes as the first such published analysis from a frontier lab. Users with 6 or more months of Claude experience showed a 3–4 percentage-point higher task success rate after controlling for task difficulty, use case, and demographic factors. Anthropic reports the effect as consistent across coding, personal tasks, and professional domains, and characterizes it as evidence of AI-specific human capital: users who invest time learning to work with AI are measurably more effective in a way separate from general skill level.
Anthropic draws several inferences from the learning-curve result. It suggests an early-adoption advantage that may compound, in which workers who adopt AI now build durable AI-collaboration skills while later adopters face a gap; it suggests that, if AI learning curves are real and predictable, employer training investments carry measurable returns; and it suggests that the 95% failure rate in enterprise AI pilots reported by the MIT NANDA GenAI Divide study may partly reflect insufficient time for users to climb the learning curve rather than integration failures alone.
Methodological limitations and caveats
The reports and surrounding analysis note several constraints:
- Self-reported success rates. Task success is measured by Claude's own assessment, a model grading its own output, which likely overstates success rates because the model does not know what the user needed.
- Selection bias. The sample comprises existing Claude users rather than a representative sample of workers or AI-naive users; early adopters likely skew toward technical, English-speaking, higher-education populations.
- Provider perspective. Anthropic has an interest in showing positive productivity impacts, and the framework is designed to highlight beneficial use such as work-related, high-complexity tasks.
- Geographic concentration. As of February 2026, the US still accounted for a dominant share of traffic, limiting global generalization.
- Short time horizon. Two data points (November 2025, February 2026) are too few to distinguish trend from noise.
Positioning in the AI adoption evidence base
The framework sits alongside other approaches to measuring AI adoption and productivity, each with different methods and constraints.
| Source | Method | Strength | Limitation | |
|---|---|---|---|---|
| [[generative-ai-at-work\ | Brynjolfsson 2023]] | RCT, customer support agents | Causal identification | Single sector, narrow task set |
| [[mit-nanda-genai-divide\ | MIT NANDA 2025]] | Survey, 1,000 enterprises | Broad organizational perspective | Self-reported; no P&L verification |
| [[metr-early-2025-dev-productivity\ | METR 2025]] | SWE tasks, controlled conditions | Rigorous task measurement | Lab conditions, not deployment |
| AEI Economic Primitives | First-party deployment logs, 2M conversations | Deployment reality at scale | Model-graded success; provider perspective |
Read together, these sources describe productivity gains that are measurable in controlled conditions while deployment reality is more variable, with benefits that require sustained user experience to accumulate.
Relationships
- supports: AI Labor Disruption — learning curves + coding dominance data
- supports: AI and Productivity — first-party deployment-reality evidence
- supports: Enterprise AI Deployment Gap — diversification and task-success data
- related: Anthropic Economic Index — January 2026: Economic Primitives — January 2026 report introducing the framework
- related: Anthropic Economic Index — March 2026: Learning Curves — March 2026 report with learning curves and three-month trends
- contradicts: MIT NANDA — The GenAI Divide (State of AI in Business 2025) — 95% pilot failure rate vs. AEI's adoption narrative (different measurement unit: AEI measures individual task success; MIT NANDA measures enterprise P&L impact)
- related: Agentic AI — coding via API migration is the clearest agentic-AI deployment signal in the data
- related: AI Deskilling — learning curves support upskilling; deskilling concerns remain for non-adopters