AI Policy Wiki
Dashboard

Sayash Kapoor

high confidence · updated 2026-08-06

PhD candidate at Princeton Computer Science; co-author with Arvind Narayanan of AI Snake Oil (2024) and 'AI as Normal Technology' (Knight Columbia, 2025). Former Facebook Core Data Science researcher. Primary junior voice (with Narayanan) for the skeptical framing of AI policy.

Sayash Kapoor is a PhD candidate in computer science at Princeton University and, with Arvind Narayanan, a co-author of the skeptical framing of AI capabilities and policy. Together they wrote the book AI Snake Oil (2024) and the essay "AI as Normal Technology" (Knight First Amendment Institute at Columbia, 2025). Before his doctoral work, Kapoor was a software engineer at Facebook Core Data Science from 2017 to 2021.

Research and writing

AI Snake Oil

AI Snake Oil (2024), co-authored with Narayanan, is a book-length treatment distinguishing predictive AI, which it characterizes as largely snake oil; generative AI, which it describes as powerful but overhyped; and content-moderation AI, which it treats as structurally hard. The argument is set out at book length in AI Snake Oil — Narayanan and Kapoor (2024).

AI as Normal Technology

"AI as Normal Technology" (2025), also co-authored with Narayanan and published by the Knight First Amendment Institute at Columbia, argues that AI is transformative but normal, that its diffusion through the economy will take decades, and that the superintelligence framing is mistaken. The academic articulation is set out in AI as Normal Technology.

Coding agents as normal technology (2026)

In "Why AI Hasn't Replaced Software Engineers, and Won't" (June 10, 2026), Kapoor and Narayanan apply the normal-technology framing to the software-engineering labor market, arguing that AI-attributed layoffs are largely "AI washing" and that software work is a "decide-execute-deliver sandwich" whose decision and delivery layers resist automation. The essay also draws on Kapoor's earlier software-engineering experience at Facebook and is the first in a planned series; it bears on AI Labor Disruption and AI Coding Agents.

Shadow evaluations of AI research agents (2026)

Kapoor conceptualized, with Narayanan, a 24-author study introducing what they call shadow evaluations, submitted to arXiv on July 29, 2026 and summarized in a companion essay on August 5 (Can AI agents conduct open-ended AI research? Early evidence from two case studies (Kirgis et al., July 2026); Source: normaltech.ai). Frontier AI agents were given the main research questions of two unpublished NeurIPS 2026 submissions, $3,000 in API credits plus GPU credits, and six days of wall-clock time; the papers' original authors then reviewed the results as conference referees. Both agent papers were unambiguously rejected, at overall scores of 2/6 and 1/6.

Reviewing over a hundred hours of agent logs, the team reported that agents quickly rejected their own proposed directions on the basis of low-quality or synthetic data; ended both runs with less than 50% of the API budget spent and hours left before the deadline, despite being able to monitor usage and being encouraged to spend down their budgets; responded to negative feedback by adding caveats and doubling down rather than shifting approach; retired their most ambitious research targets within the first day; and ignored explicit instructions on exploration time, self-review frequency and paper length. Against that, the agents completed every engineering step unaided, and the review found no significant reward hacking. The authors describe the findings as tentative, disclose a sample size of two papers across five runs, note that reviewers knew the papers were AI-generated, disclose the core team's prior position on recursive self-improvement in a dedicated section, and say they plan to increase the sample size and test new models. Peter Kirgis is first author, with Kapoor and Narayanan credited for conceptualization and Kapoor and Kirgis for drafting; the coauthors include Helen Toner, Gillian Hadfield, Seth Lazar and Rishi Bommasani.

The design is a negative-evidence complement to the autonomy measurements at METR: rather than scoring task completion, it puts the agent's output in front of the humans best placed to judge it. The under-spend and instruction-following failures bear on Agentic harnesses and capability elicitation, and the result sits alongside METR's NanoGPT-speedrun note as an empirical check on claims of automated AI research. See AI for Science, Recursive Self-Improvement (RSI).

Reproducibility crisis in machine learning

Beginning around 2022, Kapoor documented a reproducibility crisis in machine-learning-for-science papers, finding data leakage in dozens of published studies that claimed novel machine-learning capabilities for scientific prediction. This work underpins the predictive-AI portion of the AI Snake Oil argument.

COMPAS and recidivism prediction

Earlier academic work critiqued specific predictive-AI deployments in criminal justice, including the COMPAS recidivism-prediction system. This work feeds into the predictive-AI critique developed in AI Snake Oil.

Relationships