Verification asymmetry names the observation that the domains where AI systems have advanced fastest share two properties most domains lack: an output can be checked correct by a symbolic tool or an automatic verifier, and large volumes of guaranteed-correct training data can be generated cheaply. Mathematics, competitive programming, and games with formal win conditions have both. Military strategy, scientific hypothesis selection, medical judgment, and most open-ended research have neither. The argument, advanced by several sources under different phrasings, is that treating progress in the first group as an estimate of progress in the second overstates general capability.
The term describes a family of related claims rather than a single named framework; different sources emphasize the verifier, the synthetic-data supply, or the resulting evaluation problem.
The mechanism as stated
Gary Marcus gives the most compact statement, arguing after OpenAI's August 2026 mathematics announcement that "it's not an accident that math is where these models are shining. Math lends itself to two things: verification (using symbolic tools), and massive amounts of cheaply produced synthetic data where you can guarantee that the answers are correct. The same applies to coding — but it is not true in general. You can generate as many math facts as you want; you can't simulate the open-ended world. You can verify math; you can't verify a military strategy in the same way" (OpenAI's amazing — but vastly oversold — new model Astra (Marcus, August 2026)).
On this account the inference from a mathematics result to general competence is a fallacy of composition, and mathematics is a special case "which will not generalize as much people might hope." Marcus dates the argument to a point he and Ernie Davis made about Go and AlphaGo in 2019, and offers IBM's failed attempt to convert Jeopardy-winning Watson into a cancer-treatment system as the historical analogy.
Consequences for evaluation
The same asymmetry shapes what can be measured, which is where the argument bears most directly on policy. Kirgis et al. observe that the dominant evaluations of AI research-and-development automation are verifier-scored — RE-Bench, MLE-Bench, CORE-Bench, MLR-Bench, PostTrainBench — and that automatic verification "makes these evaluations objective, repeatable, and cheap to scale" while restricting them "to tasks where success can be measured as a single number." The excluded skills are named: choosing among candidate hypotheses, deciding what evidence would settle a question, and recognizing that an approach has failed and the right move is to start over.
The measurement gap becomes an inference gap. A capability curve built from verifiable benchmarks can rise steeply while capability on unverifiable work does not, and the two are not distinguishable from the benchmark alone. That study's own result is the case in point: frontier agents given six days, $3,000 in API credits and GPU credits completed every engineering step of two research projects without human help, while the original authors of the shadowed papers unambiguously rejected both agent-written papers. The paper's summary — agents "can do the engineering of AI research, but struggle with critical parts of the research lifecycle" — is the asymmetry restated as a finding.
Relation to the self-improvement debate
The asymmetry is load-bearing in disputes over Recursive Self-Improvement (RSI), because the evidence offered for AI accelerating AI research is drawn disproportionately from the verifiable side. Kirgis et al. note that Anthropic's "When AI Builds Itself" "explicitly cites Claude's rising success rate on LLM-judged 'open-ended' Claude Code sessions as evidence of self-improvement," and that an Elasticity Institute report on the economics of recursive self-improvement distinguishes "broad" from "narrow" AI capabilities and calls for more data on the full breadth of capabilities relative to human researchers. Anthropic's own essay locates the remaining gap in the same place, describing Claude as at or above human level at execution while trailing at direction-setting and research taste.
The counter-position is that full automation of unverifiable work may not be on the critical path. Kirgis et al. state it themselves: it is "possible that the path to AI R&D does not require full automation of open-ended tasks like the ones we study, or that the open-ended research skills that we measure are not ones on the critical path." Neither side has published a measure that would settle which skills are load-bearing.
Open questions
- Whether the boundary is fixed or moving. Formal verification tooling extends to new domains over time, and autoformalization would widen the verifiable set substantially; Davis argues that problem "does not seem close to being solved" (OpenAI's amazing — but vastly oversold — new model Astra (Marcus, August 2026)).
- What fraction of frontier AI development consists of open-ended research rather than hill-climbing on well-specified objectives. Kirgis et al. name this as something they "do not know."
- Whether expert-graded open-world evaluations can be made repeatable enough to track a trend rather than supply single datapoints. See Shadow Evaluations.
Relationships
- supports: Jagged Frontier — the uneven capability profile this argument predicts
- related: Shadow Evaluations — a method built specifically to evaluate the unverifiable side
- related: Recursive Self-Improvement (RSI) — the forecasting debate the asymmetry bears on
- related: AI Benchmarks and Evaluation, AI for Science, Reward Hacking
- related: OpenAI's amazing — but vastly oversold — new model Astra (Marcus, August 2026), Can AI agents conduct open-ended AI research? Early evidence from two case studies (Kirgis et al., July 2026), When AI Builds Itself (The Anthropic Institute, June 4, 2026)