AI Policy Wiki
Dashboard

OpenAI's amazing — but vastly oversold — new model Astra (Marcus, August 2026)

medium confidence · updated 2026-08-06

Gary Marcus's August 2, 2026 essay on the reaction to OpenAI's claim that an internal version of Astra solved ten open problems in mathematics and theoretical computer science. Argues the reaction commits the fallacy of composition, and gives a principled reason to expect mathematics not to generalize: math permits symbolic verification and cheap generation of guaranteed-correct synthetic data, neither of which is available for problems that are hard to formalize. Carries Ernie Davis's evaluation critique on the missing denominator and the true cost of the result.

An essay published August 2, 2026 by Gary Marcus on his newsletter Marcus on AI, and republished in full by the ACM under BLOG@CACM. It responds to the reaction to OpenAI's August 1 announcement that an internal version of the unreleased Astra family had solved ten open problems in mathematics, quantum complexity and theoretical computer science (Ten Advances in Mathematics and Theoretical Computer Science). The essay is structured in four parts and is a position document rather than an empirical report; its distinctive contributions are the fallacy-of-composition framing, the verification-and-synthetic-data account of why mathematics is a special case, and a long quoted critique from Ernie Davis.

Marcus does not dispute the result. The essay opens by conceding that Astra "is amazing. No denying that," and repeats the concession at the close of each argument. What it disputes is the inference drawn from it.

The fallacy of composition

Marcus's central claim is that a set of widely circulated reactions all commit the same error — treating success at one kind of cognition as evidence that success at all kinds is imminent. He names it the fallacy of composition and breaks it into two steps: "Someone pretends that all cognition is created equally," and then "Whenever AI achieves success on some form of (fancy) cognition, they want you to believe that success on all forms of AI is imminent."

The essay quotes four reactions as instances. Dean Ball wrote that "everyone in the world will soon be able to use the model that made these breakthroughs for every problem they face in life, no matter how mundane." Matt Shumer wrote that the next OpenAI model would "make Fable look like a toy, and usher in a golden age of science." A further post claimed "the species just crossed a one-way threshold," and Elon Musk treated the announcement as evidence of the Singularity.

The counter-argument is drawn from the structure of human expertise: "expertise in math doesn't guarantee genius in all domains," which Marcus connects to the multidimensional theories of intelligence of Howard Gardner and Robert Sternberg and to the separation of math and verbal sections on the SAT. Applied to Astra, the conclusion is that competence at some mathematics does not imply the model "will avoid hallucinations or solve the reliability problems other GenAI systems have," will read PDFs reliably, or will be "the first generative AI to be able to obey hard rules." Marcus states he sees "no reason whatsoever to think Astra is AGI let alone ASI," and would be surprised if it scored 5/10 on his 2024 bet with Miles Brundage.

Why mathematics is argued to be a special case

The second part supplies the mechanism Marcus says makes the fallacy applicable here rather than merely possible. Mathematics, on his account, has two properties that most domains lack: it permits verification using symbolic tools, and it permits "massive amounts of cheaply produced synthetic data where you can guarantee that the answers are correct." Coding shares both. Open-ended domains share neither: "You can generate as many math facts as you want; you can't simulate the open-ended world. You can verify math; you can't verify a military strategy in the same way."

Marcus dates the argument to at least January 2025 and traces it to a point he and Ernie Davis made about Go and AlphaGo in 2019. The historical analogy offered is IBM's attempt to convert Jeopardy-winning Watson into a cancer-treatment system, which he describes as having failed. See AI for Science and Verification Asymmetry.

Davis's evaluation critique

Ernie Davis's contribution, quoted at length from an email on the essay's first draft, is the element of the source that turns on evaluation methodology rather than inference.

The denominator. Davis asks how many conjectures were attempted, and sets out how differently the result reads under three scenarios: ten solved out of ten drawn at random from all outstanding significant conjectures "would be amazing"; ten out of a cherry-picked fifty is "still amazing, but significantly less so"; and a run across "all 1000 or so open Erdos conjectures and on 10,000 other open conjectures" would make the failures "significant information on its limits as a mathematical reasoner."

The cost. OpenAI's roughly $2,000 figure, Davis argues, is "a safe bet" to cover only the conjectures where Astra succeeded, and excludes "the cost in terms of the salaries of the highly-paid mathematicians and computer scientists who worked on this project," which he estimates at "not less than $20,000" and possibly "upward of $200,000." He adds that comparable information has never been released for the system that reached gold-medal performance at the 2025 International Mathematical Olympiad.

Historical scale. Against the claim that the debut was "plausibly the most significant day in the history of mathematics," Davis observes that 14 of David Hilbert's 23 problems have been solved since 1900 — "one every nine years — and these results are nowhere near that league," adding that "you could easily compile a list of 100 much more important results that have been proved since 1926."

Autoformalization. Davis notes that turning human-written mathematics into strictly logical form "does not seem close to being solved," citing Kevin Buzzard's multi-year Lean formalization of Wiles's proof of Fermat's Last Theorem as a task no current system can carry out: "Presumably if they could do it unassisted, they would scoop him, and if they could cut down his work from years to weeks, he would be using them."

Further points

Marcus's remaining points concern disclosure and scope. He characterizes the announcement as "marketing, not science," noting that neither the post nor the accompanying 249-page article says "how the model works, how the proofs were verified, what role if any humans played, whether any of the proposed proofs had errors." He cites Eric Weinstein on kinds of mathematics the model may handle less well, and Henry Yuen's observation that the write-up of one proof reads like model-generated prose, elaborating on boilerplate setup and then introducing key steps without support — a gap between proof quality and exposition quality. He doubts the result will produce a cure for cancer or advances across "materials research, energy production, drug discovery," and rejects the claim that it marks the arrival of "the era of automated scientific discovery," pointing to a then-circulating MIT and Harvard paper, "Evaluating Large Language Models in Scientific Discovery," and to recent work by Arvind Narayanan and coauthors.

The essay records that the pushback Marcus received on X was "most of it ad hominem, some literally involving fabrication."

Relationships