HealthBench Professional Leaderboard
Last updated September 30, 2026 · 31 models ranked
GPT-6 Astra (Anthropic run) leads the HealthBench Professional leaderboard at 0.703, ahead of Claude Sonnet 5.5 (0.692) and Claude Fable 5 (0.660). The benchmark, published by OpenAI in April 2026, scores models on 525 tasks physicians selected from real clinician conversations, graded against physician-written rubrics on a 0 to 1 scale. Physician-written responses score 0.437 on the same rubrics.
Leaderboard
Sources525-task set, scores compiled from published documents. Updated September 30, 2026.
GPT-6 Astra (Anthropic run)
Claude Sonnet 5.5
Claude Fable 5
Claude Opus 5.5
GPT-6 Astra
Claude Fable 5 (September card)
Claude Fable 5.1
GPT-6 Luna
GPT-6 Sol
GPT-5.6 Sol
Claude Opus 5
Muse Spark 1.1
Claude Sonnet 5
GPT-5.6 Terra
Claude Opus 4.8 (Opus 4.8 grader)
Grok 4.7
Claude Opus 4.8
GPT-5.6 Luna
Muse Spark
GPT-5.6 Sol (August)
Claude Opus 4.7
GPT-5.5
Grok 4.6
GPT-5.4
GPT-5
GPT-5.2
Claude Sonnet 4.6
GPT-5.6 Luna (August)
GPT-5.1
GPT-5.5 Instant
MAI-Thinking-1
physician baseline 0.437
Full ranking
Sources| # | model | score | size | context | cost in / out per 1M | license | |
|---|---|---|---|---|---|---|---|
| 1 | GPT-6 Astra (Anthropic run) OpenAI | 0.703 | — | 1.05M | $10.00 / $50.00 | proprietary | |
| 2 | Claude Sonnet 5.5 Anthropic | 0.692 | — | 1M | $2.00 / $10.00 | proprietary | |
| 3 | Claude Fable 5 Anthropic | 0.660 | — | 1.0M | $10.00 / $50.00 | proprietary | |
| 4 | Claude Opus 5.5 Anthropic | 0.656 | — | 1M | $4.00 / $20.00 | proprietary | |
| 5 | GPT-6 Astra OpenAI | 0.647 | — | 1.05M | $10.00 / $50.00 | proprietary | |
| 6 | Claude Fable 5 (September card) Anthropic | 0.633 | — | 1.0M | $10.00 / $50.00 | proprietary | |
| 7 | Claude Fable 5.1 Anthropic | 0.621 | — | 1.0M | $10.00 / $50.00 | proprietary | |
| 8 | GPT-6 Luna OpenAI | 0.608 | — | 1.05M | $0.10 / $0.50 | proprietary | |
| 9 | GPT-6 Sol OpenAI | 0.608 | — | 1.05M | $2.00 / $10.00 | proprietary | |
| 10 | GPT-5.6 Sol OpenAI | 0.605 | — | 1.05M | $4.00 / $20.00 | proprietary | |
| 11 | Claude Opus 5 Anthropic | 0.598 | — | 1.0M | $5.00 / $25.00 | proprietary | |
| 12 | Muse Spark 1.1 Meta | 0.593 | — | 1.0M | $1.25 / $4.25 | proprietary | |
| 13 | Claude Sonnet 5 Anthropic | 0.578 | — | 1.0M | $2.00 / $10.00 | proprietary | |
| 14 | GPT-5.6 Terra OpenAI | 0.577 | — | 1.05M | $2.00 / $12.00 | proprietary | |
| 15 | Claude Opus 4.8 (Opus 4.8 grader) Anthropic | 0.574 | — | 1.0M | $5.00 / $25.00 | proprietary | |
| 16 | Grok 4.7 SpaceX AI | 0.567 | — | 500K | $2.00 / $6.00 | proprietary | |
| 17 | Claude Opus 4.8 Anthropic | 0.558 | — | 1.0M | $5.00 / $25.00 | proprietary | |
| 18 | GPT-5.6 Luna OpenAI | 0.557 | — | 1.05M | $0.20 / $1.20 | proprietary | |
| 19 | Muse Spark Meta | 0.541 | — | 1.0M | — | proprietary | |
| 20 | GPT-5.6 Sol (August) OpenAI | 0.540 | — | 1.05M | $4.00 / $20.00 | proprietary | |
| 21 | Claude Opus 4.7 Anthropic | 0.519 | — | 1M | $5.00 / $25.00 | proprietary | |
| 22 | GPT-5.5 OpenAI | 0.518 | — | 1.05M | $5.00 / $30.00 | proprietary | |
| 23 | Grok 4.6 SpaceX AI | 0.485 | — | 500K | $2.00 / $6.00 | proprietary | |
| 24 | GPT-5.4 OpenAI | 0.481 | — | 1.05M | $2.50 / $15.00 | proprietary | |
| 25 | GPT-5 OpenAI | 0.462 | — | 400K | $1.25 / $10.00 | proprietary | |
| 26 | GPT-5.2 OpenAI | 0.459 | — | 400K | $1.75 / $14.00 | proprietary | |
| 27 | Claude Sonnet 4.6 Anthropic | 0.442 | — | 1M | $3.00 / $15.00 | proprietary | |
| 28 | GPT-5.6 Luna (August) OpenAI | 0.441 | — | 1.05M | $0.20 / $1.20 | proprietary | |
| 29 | GPT-5.1 OpenAI | 0.396 | — | 400K | $1.25 / $10.00 | proprietary | |
| 30 | GPT-5.5 Instant OpenAI | 0.384 | — | 400K | $5.00 / $30.00 | proprietary | |
| 31 | MAI-Thinking-1 Microsoft | 0.350 | 1.0T | 256K | $2.00 / $8.00 | proprietary | |
Each score is copied as its source document prints it, on the 525-task set. The line under each model name says which document the score was read from and links to its entry on the sources page. List API prices per 1M tokens as of September 30, 2026. GPT-5.6 models charge higher rates above 272K input tokens. MAI-Thinking-1 is in public preview on Microsoft Foundry without final list pricing. Each model name links to its full result.
What the benchmark measures
HealthBench Professional evaluates language models on the work clinicians actually bring to a model: care consults such as differential diagnosis and management questions, clinical writing and documentation, and medical research. The 525 tasks were selected by physicians from 15,079 real clinician conversations, difficult cases were deliberately overweighted, and roughly a third of the set is adversarial. 190 physicians across 50 countries and 26 specialties wrote and adjudicated the rubrics. The full construction is described on the benchmark page.
How to read the scores
Scores are earned rubric points divided by possible points, clipped to 0 to 1, with a length adjustment that keeps verbosity from buying credit. Absolute numbers run low by design: the field's best result is 0.703, and physician-written responses score 0.437 on the same rubrics. Scores depend on the grader version, the reasoning effort, and the deployment setting a document reports, so 2 numbers taken from different documents can differ for reasons other than model quality. The scores here are compiled from published sources, not produced by Arcophos; where each one comes from and how it is read is on the methodology page.
Head to head
Score, price, and context differences for the pairings people actually weigh, with the full score-difference matrix on the compare page.
- 0.660 vs 0.605 · Claude Fable 5 by 0.055
- 0.605 vs 0.598 · GPT-5.6 Sol by 0.007
- 0.578 vs 0.577 · Claude Sonnet 5 by 0.001
- 0.660 vs 0.598 · Claude Fable 5 by 0.062
- 0.598 vs 0.558 · Claude Opus 5 by 0.040
- 0.557 vs 0.384 · GPT-5.6 Luna by 0.173
Data
The full results are published as JSON and CSV at stable URLs, licensed CC BY 4.0. Citation format and version history are on the data page.