HealthBench is OpenAI's evaluation framework for health-related conversations, developed alongside ChatGPT Health and used as the reported measure of that product's performance.
What it measures
The design choice that distinguishes HealthBench from earlier medical AI benchmarks is what it declines to measure. Rather than exam-style accuracy — the USMLE-question format on which language models had already reached high scores — it uses physician-written rubrics scoring safety, clarity, escalation appropriateness, and respect for individual context (Source: openai.com).
Escalation appropriateness is the most consequential of the four for consumer deployment: it scores whether the system directs a user to clinical care when it should, which is a property of the interaction rather than of the model's knowledge. A system can answer a question correctly and still fail by answering it at all.
Construction
The evaluation was built from feedback by "260 or more physicians across 60 countries and 30 specialties providing feedback more than 600,000 times over two years" (Source: openai.com). The breadth across countries and specialties is a response to a standing criticism of medical AI benchmarks — that they encode the norms of a single health system — though the resulting rubrics are not published in a form that permits independent replication.
Use and reported results
OpenAI reports HealthBench scores as the performance measure for ChatGPT Health, including gains for GPT-5.6-Luna at the July 23, 2026 general-availability launch (Source: techcrunch.com). Meta also reports HealthBench results — across HealthBench, HealthBench Hard, and HealthBench Consensus variants — in its Muse Spark report, indicating adoption as a cross-developer comparison point rather than an OpenAI-internal measure.
The evidence gap it does not close
HealthBench measures conversational conduct as judged by physicians against rubrics. It does not measure whether users who receive those conversations experience better health outcomes. The ChatGPT Health launch page "cites no peer-reviewed clinical-outcome studies," a gap Topol/Marcus: LLMs and Patient Outcomes — three-document cluster addresses in finding "very little evidence for LLMs benefiting patients or doctors for health outcomes" outside administrative work (Source: openai.com). The distinction matters for how HealthBench results should be read in policy contexts: a rubric score is evidence about the interaction, not about the endpoint.
The gap is also what keeps the product on one side of a regulatory line. OpenAI's framing that ChatGPT Health is "designed to support, not replace" clinicians and "not intended for diagnosis or treatment" maintains distance from software-as-a-medical-device classification, under which outcome evidence would be required (Source: openai.com). See Healthcare — AI Deployment.
Relationships
- related: Healthcare — AI Deployment — the sector context and the ROI-evidence debate
- related: AI Benchmarks and Evaluation — a rubric-based alternative to exam-style medical benchmarks
- related: OpenAI, Topol/Marcus: LLMs and Patient Outcomes — three-document cluster, Muse Spark Safety & Preparedness Report (Meta, May 2026), AI Deskilling