GPQA is a dataset of 448 graduate-level multiple-choice questions in biology, physics, and chemistry, written and peer-validated by domain experts. The paper introducing it was published by David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman (NYU) as arXiv 2311.12022 in November 2023.
Summary of findings
The dataset is described as "Google-proof": non-expert validators given roughly 30 minutes or more of web search reach only 34% accuracy, compared with 65% for PhD domain experts. The authors attribute this gap to question design that requires graduate-level synthesis rather than retrieval. The stated motivation is scalable oversight: GPQA is built to test regimes where even careful non-experts struggle to verify AI outputs, serving as a proxy for hard-to-verify domains. Bowman's group framed the benchmark as infrastructure for scalable-oversight experiments.
Key claims
- PhD experts reach 65% accuracy, non-expert web-searchers 34%, and GPT-4 (late 2023) 39% at the time of release. (high)
- Difficulty derives from question design that requires graduate-level synthesis, not retrieval. (high)
- The benchmark functions as a proxy for hard-to-verify domains and was designed to support scalable-oversight research. (high)
Reception and use
GPQA is cited in frontier-model system cards including Claude Opus 4.6 System Card, Claude Sonnet 4.6 System Card, GPT-5.3-Codex System Card, and GPT-5.4 Thinking System Card. It is a reference point for the AI Benchmarks and Evaluation concept and for tracking reasoning-capability saturation: frontier models now exceed 80% on GPQA, above the 65% PhD-expert baseline, within about two years of the benchmark's release.
Relationships
- supports: AI Benchmarks and Evaluation, AI Software Progress
- related: Safety Cases for Frontier AI, AI Scheming (scalable oversight framing)
- instance-of: AI Benchmarks and Evaluation