HealthBench Professional

FAQ

Short answers to the questions the leaderboard gets asked. Each answer stands on its own; the pages linked at the bottom carry the detail.

Which AI model scores highest on HealthBench Professional?

Claude Fable 5 scores highest at 0.660 as of August 16, 2026, ahead of GPT-5.6 Sol at 0.605 and Claude Opus 5 at 0.598. 9 models are on the leaderboard.

What is HealthBench Professional?

HealthBench Professional is a benchmark published by OpenAI in April 2026. It contains 525 physician-authored tasks selected from 15,079 real clinician conversations, covering care consults, clinical writing and documentation, and medical research. Each response is graded against a physician-written rubric.

How is HealthBench Professional different from HealthBench?

The original HealthBench, from May 2025, covers 5,000 general health conversations with laypeople and professionals. HealthBench Professional narrows the scope to what clinicians bring to a model at work, uses tasks drawn from real clinician conversations, and enriches for difficulty, with roughly a third of the set written adversarially.

How are the scores graded?

A model grader (GPT-5.4, low reasoning effort) judges each rubric criterion as met or not met. The score is earned points divided by possible points, clipped to 0 to 1, with a length adjustment so verbose answers cannot buy credit. Rubrics include negative criteria that subtract points for harmful or fabricated content.

Do AI models beat physicians on this benchmark?

On these rubrics, yes: physician-written responses score 0.437, and 7 of the 9 models score above that. It is a statement about rubric adherence on 525 hard, written tasks, not about clinical judgment, procedures, or accountability, so it should be read narrowly.

Which model gives the most score per dollar?

GPT-5.6 Luna, at 0.557 for $0.20 per million input tokens and $1.20 per million output tokens, scores within 0.103 of the leader at a fraction of the price of every model above it.

How often is the leaderboard updated?

The leaderboard is re-run when a frontier model ships or is materially updated. The current table covers 9 models as of August 16, 2026, and every change is dated on the updates page.

Can I download the results?

Yes. The full results are published as JSON and CSV at stable URLs under a CC BY 4.0 license, with a suggested citation format on the data page.

The full ranking is on the leaderboard. The benchmark's construction is on the benchmark page, the run protocol on the methodology page, and downloads on the data page.