AI Policy Wiki
Dashboard

MIT NANDA — The GenAI Divide (State of AI in Business 2025)

high confidence · updated 2026-06-06

MIT Project NANDA report finding 95% of $30-40B enterprise GenAI pilots yield no measurable P&L impact; identifies the 'learning gap' as the core barrier and finds vendor partnerships ~2x more successful than internal builds.

The GenAI Divide is a July 2025 report from MIT Media Lab's Project NANDA (Networked Agents and Decentralized AI) on the state of enterprise generative-AI adoption. Drawing on a review of 300+ public enterprise AI initiatives, 52 structured executive interviews, and a survey of 153 senior leaders, it reports that after an estimated $30–40 billion in enterprise GenAI spend, roughly 95% of organizations see no measurable profit-and-loss impact, and that only about 5% of custom enterprise AI tools reach production. The report attributes the shortfall to what it terms a "learning gap" rather than to model quality, compute, talent, or regulation.

Authors: Aditya Challapally (lead), Chris Pease, Ramesh Raskar, Pradyumna Chari Institution: MIT Media Lab — Project NANDA (Networked Agents and Decentralized AI) Published: July 2025 Raw source: Raw Sources/MIT NANDA - The GenAI Divide

Summary of argument and findings

The report's central empirical claim is that enterprise GenAI deployment is underperforming relative to headline capability gains. Its headline figures are that after $30–40 billion in enterprise spend, roughly 95% of organizations report no measurable P&L impact, and only about 5% of custom enterprise AI tools reach production.

Rather than attribute the shortfall to the usual candidate explanations — model quality, compute, talent, or regulation — the report locates the binding constraint in a "learning gap": most deployed GenAI tools are stateless, do not retain feedback, and cannot adapt to firm-specific workflows. On this framing, the enterprise AI problem is recast from one of "better models" to one of building systems with memory that integrate deeply into workflows.

The report also reports a difference between deployment paths: the vendor-partnership route succeeds about 67% of the time, against about 33% for internal builds, roughly twice the success rate. It finds that back-office automation yields higher return on investment than sales and marketing applications, despite more than 50% of budget going to the latter. It describes a "shadow AI economy" in which more than 90% of employees use personal LLM accounts for work. On workforce effects, it characterizes the impact as position attrition concentrated in previously-outsourced roles rather than layoffs at the largest firms, and it reports structural disruption in the technology and media sectors while finance, healthcare, and energy remain largely unchanged.

Key claims

The table records the report's principal claims with a confidence assessment for each.

#ClaimConfidenceNotes
1~95% of enterprise GenAI pilots show no measurable P&L impactHighHeadline finding, widely cited; survey + public-initiative methodology
2$30–40B estimated total enterprise GenAI spend to mid-2025MediumOrder-of-magnitude estimate, plausible given other analyst figures
3Vendor-partnership path succeeds ~67% vs. ~33% for internal buildsHighConsistent across multiple secondary summaries
4Core barrier is "learning gap" (no feedback retention, no context memory, no workflow adaptation)High (as the report's framing); Medium (as the causal claim)The finding that enterprises want these features is high-confidence survey data; the causal claim that absence causes the 95% failure is an interpretation
5Back-office automation yields higher ROI than sales/marketing despite >50% of budget going to the latterHighQuantitative finding from the survey
6>90% of employees use personal LLM accounts for work ("shadow AI economy")MediumSelf-reported; number consistent across coverage
7Workforce impact is position attrition concentrated in previously-outsourced roles, not layoffs at the S&P topMediumFraming claim; harder to verify from the survey alone
8Tech/media sectors show structural disruption; finance/healthcare/energy largely unchangedMediumSectoral generalization from a modest sample

Methodology

The findings rest on three inputs: a desk review of 300+ publicly disclosed enterprise AI initiatives, 52 structured executive interviews, and 153 senior-leader survey responses. Fortune's reporting references "150 interviews / 350 employees," which likely conflates the structured-executive sample with a wider employee survey; exact reconciliation depends on the final report deck.

Interpretation and limitations

Several of the report's figures carry caveats. The $30–40 billion figure is an order-of-magnitude estimate and is best treated as directional rather than audited. The 95% number is a share-of-initiatives figure, not a share-of-dollars: a minority of successful deployments could still absorb a majority of spending. The vendor-partnership versus internal-build comparison is presented as a correlation rather than a causal result; firms that choose vendors may differ systematically from firms that choose to build, for example in having more focused use cases or better-defined workflows, and a randomized comparison is not possible. The "learning gap" frame is also central to Project NANDA's broader research agenda on agentic, memory-enabled systems, so the report functions simultaneously as an empirical diagnosis and as a case for the kind of system NANDA is building toward.

Reception and relation to other work

The report's enterprise-portfolio findings sit alongside, rather than against, task-level productivity studies. Stanford HAI's ethnography by Karunakaran, Vendraminelli, and Narayanan of a four-year fashion-company deployment found a similar pattern at micro scale, turning on jurisdictional clarity, task centrality, and task-enactment homogeneity; MIT NANDA generalizes that pattern across 300+ initiatives, and the two are mutually reinforcing (Source: hai.stanford.edu). The slow-deployment, organizational-bottleneck account is consistent with AI Diffusion, which provides direct evidence that AI diffusion within organizations is slower than headline capability would suggest, and with the AI as Normal Technology framing.

The report's deployment-level pessimism contrasts with, but does not contradict, task-level productivity findings. Generative AI at Work (Brynjolfsson, Li, Raymond 2023) reported a +15% task-level productivity gain in customer support, and Stanford HAI AI Index Report 2026 reported 14–26% productivity gains in customer support and software development. Both can hold simultaneously with NANDA's finding: AI raises measured productivity on specific tasks where it is successfully deployed, while the majority of enterprise deployments never reach that stage. The gap is between demonstrated capability and realized deployment. On this reading, NANDA measures the share of attempts that become successful deployments, while the task-level studies measure the gains those successful deployments deliver. The report adds an enterprise-portfolio counterweight to the task-level studies in AI and Productivity.

The report's proposed direction — memory, feedback retention, and workflow adaptation — maps onto the Agentic AI capability envelope, making the report implicitly an argument for why agentic systems matter commercially and not only technically. It is an instance of the broader Enterprise AI Deployment Gap pattern.

No contradictions with existing wiki claims were identified on review. The report complements task-level productivity findings with enterprise-portfolio-level deployment findings. A page implying that enterprise GenAI broadly delivers ROI on the scale implied by headline spend figures would need revision against this source, but none was found.

Relationships