FAQ
Short answers to the questions the leaderboard gets asked. Each answer stands on its own; the pages linked at the bottom carry the detail.
Which AI model scores highest on HealthBench Professional?
GPT-6 Astra (Anthropic run) scores highest at 0.703 as of September 30, 2026, ahead of Claude Sonnet 5.5 at 0.692 and Claude Fable 5 at 0.660. 31 models are on the leaderboard.
What is HealthBench Professional?
HealthBench Professional is a benchmark published by OpenAI in April 2026. It contains 525 physician-authored tasks selected from 15,079 real clinician conversations, covering care consults, clinical writing and documentation, and medical research. Each response is graded against a physician-written rubric.
How is HealthBench Professional different from HealthBench?
The original HealthBench, from May 2025, covers 5,000 general health conversations with laypeople and professionals. HealthBench Professional narrows the scope to what clinicians bring to a model at work, uses tasks drawn from real clinician conversations, and enriches for difficulty, with roughly a third of the set written adversarially.
How are the scores graded?
A model grader (GPT-5.4, low reasoning effort) judges each rubric criterion as met or not met. The score is earned points divided by possible points, clipped to 0 to 1, with a length adjustment so verbose answers cannot buy credit. Rubrics include negative criteria that subtract points for harmful or fabricated content.
Do AI models beat physicians on this benchmark?
On these rubrics, yes: physician-written responses score 0.437, and 28 of the 31 models score above that. It is a statement about rubric adherence on 525 hard, written tasks, not about clinical judgment, procedures, or accountability, so it should be read narrowly.
Which model gives the most score per dollar?
GPT-5.6 Luna, at 0.557 for $0.20 per million input tokens and $1.20 per million output tokens, scores within 0.146 of the leader at a fraction of the price of every model above it.
Who ran these evaluations?
Not Arcophos. The scores on this site are compiled from published documents: OpenAI system cards and update reports, Anthropic system cards, other vendors' model cards and announcements, and named leaderboards. Each row links to the document its number was read from on the sources page, and rows with no located document are marked source pending. 0 of 31 rows are source pending as of September 30, 2026.
Why does a model's score here differ from the one in its system card or in the paper?
Because the 2 numbers were produced under different settings. HealthBench Professional scores move with the length adjustment, the reasoning effort the model was run at, the deployment setting (a product such as ChatGPT for Clinicians against the base model), and the grader version. The sources page shows the exact excerpt each number was read from so the settings can be compared.
How often is the leaderboard updated?
The board changes when a vendor publishes a new HealthBench Professional result, when a document is located for a row that was source pending, or when a higher-authority document replaces an existing source. The current table covers 31 models as of September 30, 2026, and every change is dated on the updates page.
Can I download the results?
Yes. The full results are published as JSON and CSV at stable URLs under a CC BY 4.0 license, with a suggested citation format on the data page.
The full ranking is on the leaderboard. The benchmark's construction is on the benchmark page, where each score comes from on the methodology page and the sources page, and downloads on the data page.