AI Policy Wiki
Dashboard

Ernie Davis

medium confidence · updated 2026-08-06

Professor of computer science at New York University and long-time collaborator of Gary Marcus; a recurring source of evaluation-methodology critiques of AI capability claims, particularly on missing denominators, true costs, and commonsense reasoning benchmarks.

Ernest Davis is a professor of computer science at New York University and a frequent collaborator of Gary Marcus, with whom he co-authored Rebooting AI (2019). His recurring contribution to the AI capability debate is methodological: given a claimed result, he asks what would have to be disclosed for the claim to be assessable, and identifies which of those items is missing.

Evaluation critiques

The denominator problem

Davis's most-cited objection concerns the ratio a headline result is drawn from. Quoted at length in Marcus's August 2, 2026 essay on OpenAI's claim that an internal version of Astra had solved ten open problems in mathematics and theoretical computer science, Davis set out how differently the same ten results read under three scenarios: ten solved from ten drawn at random from all outstanding significant conjectures "would be amazing"; ten from a cherry-picked fifty is "still amazing, but significantly less so"; and a run across "all 1000 or so open Erdos conjectures and on 10,000 other open conjectures" would make the failures "significant information on its limits as a mathematical reasoner." OpenAI did not state how many conjectures were attempted (OpenAI's amazing — but vastly oversold — new model Astra (Marcus, August 2026)).

The cost denominator

On the same result, Davis argued that OpenAI's roughly $2,000 figure was "a safe bet" to cover only successful runs, and that the material omission was human labour: "what was the cost in terms of the salaries of the highly-paid mathematicians and computer scientists who worked on this project? I'd be astonished if it was less than $20,000 and would not be surprised if it was upward of $200,000." He added that comparable information had never been released for the system that achieved gold-medal performance at the 2025 International Mathematical Olympiad (OpenAI's amazing — but vastly oversold — new model Astra (Marcus, August 2026)).

Historical calibration

Against the claim that the announcement was "plausibly the most significant day in the history of mathematics," Davis offered a base rate: 14 of David Hilbert's 23 problems have been solved since 1900, "one every nine years — and these results are nowhere near that league," adding that "you could easily compile a list of 100 much more important results that have been proved since 1926" (OpenAI's amazing — but vastly oversold — new model Astra (Marcus, August 2026)).

Autoformalization

Davis has argued that converting human-written mathematics into strictly logical form "does not seem close to being solved," using Kevin Buzzard's multi-year Lean formalization of Wiles's proof of Fermat's Last Theorem as the test case: "It is quite safe to say that no AI system is capable of carrying out that monumental task. Presumably if they could do it unassisted, they would scoop him, and if they could cut down his work from years to weeks, he would be using them" (OpenAI's amazing — but vastly oversold — new model Astra (Marcus, August 2026)).

Earlier work with Marcus

Davis and Marcus argued in 2019, in the context of AlphaGo, that success in a closed, formally scored domain would not generalize to open-ended ones — the argument Marcus later applied to mathematics and reinforcement learning on verifiable rewards (see Verification Asymmetry). Marcus has also credited Davis as co-author of an earlier statement of the proposition that a model's proof-writing quality can lag the quality of its proofs.

Relationships