HealthBench Professional

HealthBench Professional Leaderboard

Last updated September 30, 2026 · 31 models ranked

GPT-6 Astra (Anthropic run) leads the HealthBench Professional leaderboard at 0.703, ahead of Claude Sonnet 5.5 (0.692) and Claude Fable 5 (0.660). The benchmark, published by OpenAI in April 2026, scores models on 525 tasks physicians selected from real clinician conversations, graded against physician-written rubrics on a 0 to 1 scale. Physician-written responses score 0.437 on the same rubrics.

Leaderboard

Sources

525-task set, scores compiled from published documents. Updated September 30, 2026.

  1. 0.703
    OpenAI logoGPT-6 Astra (Anthropic run)
  2. 0.692
    Anthropic logoClaude Sonnet 5.5
  3. 0.660
    Anthropic logoClaude Fable 5
  4. 0.656
    Anthropic logoClaude Opus 5.5
  5. 0.647
    OpenAI logoGPT-6 Astra
  6. 0.633
    Anthropic logoClaude Fable 5 (September card)
  7. 0.621
    Anthropic logoClaude Fable 5.1
  8. 0.608
    OpenAI logoGPT-6 Luna
  9. 0.608
    OpenAI logoGPT-6 Sol
  10. 0.605
    OpenAI logoGPT-5.6 Sol
  11. 0.598
    Anthropic logoClaude Opus 5
  12. 0.593
    Meta logoMuse Spark 1.1
  13. 0.578
    Anthropic logoClaude Sonnet 5
  14. 0.577
    OpenAI logoGPT-5.6 Terra
  15. 0.574
    Anthropic logoClaude Opus 4.8 (Opus 4.8 grader)
  16. 0.567
    SpaceX AI logoGrok 4.7
  17. 0.558
    Anthropic logoClaude Opus 4.8
  18. 0.557
    OpenAI logoGPT-5.6 Luna
  19. 0.541
    Meta logoMuse Spark
  20. 0.540
    OpenAI logoGPT-5.6 Sol (August)
  21. 0.519
    Anthropic logoClaude Opus 4.7
  22. 0.518
    OpenAI logoGPT-5.5
  23. 0.485
    SpaceX AI logoGrok 4.6
  24. 0.481
    OpenAI logoGPT-5.4
  25. 0.462
    OpenAI logoGPT-5
  26. 0.459
    OpenAI logoGPT-5.2
  27. 0.442
    Anthropic logoClaude Sonnet 4.6
  28. 0.441
    OpenAI logoGPT-5.6 Luna (August)
  29. 0.396
    OpenAI logoGPT-5.1
  30. 0.384
    OpenAI logoGPT-5.5 Instant
  31. 0.350
    Microsoft logoMAI-Thinking-1

physician baseline 0.437

Full ranking

Sources
#modelscoresizecontextcost in / out per 1Mlicense
1OpenAI logoGPT-6 Astra (Anthropic run) OpenAI0.703—1.05M$10.00 / $50.00proprietary
2Anthropic logoClaude Sonnet 5.5 Anthropic0.692—1M$2.00 / $10.00proprietary
3Anthropic logoClaude Fable 5 Anthropic0.660—1.0M$10.00 / $50.00proprietary
4Anthropic logoClaude Opus 5.5 Anthropic0.656—1M$4.00 / $20.00proprietary
5OpenAI logoGPT-6 Astra OpenAI0.647—1.05M$10.00 / $50.00proprietary
6Anthropic logoClaude Fable 5 (September card) Anthropic0.633—1.0M$10.00 / $50.00proprietary
7Anthropic logoClaude Fable 5.1 Anthropic0.621—1.0M$10.00 / $50.00proprietary
8OpenAI logoGPT-6 Luna OpenAI0.608—1.05M$0.10 / $0.50proprietary
9OpenAI logoGPT-6 Sol OpenAI0.608—1.05M$2.00 / $10.00proprietary
10OpenAI logoGPT-5.6 Sol OpenAI0.605—1.05M$4.00 / $20.00proprietary
11Anthropic logoClaude Opus 5 Anthropic0.598—1.0M$5.00 / $25.00proprietary
12Meta logoMuse Spark 1.1 Meta0.593—1.0M$1.25 / $4.25proprietary
13Anthropic logoClaude Sonnet 5 Anthropic0.578—1.0M$2.00 / $10.00proprietary
14OpenAI logoGPT-5.6 Terra OpenAI0.577—1.05M$2.00 / $12.00proprietary
15Anthropic logoClaude Opus 4.8 (Opus 4.8 grader) Anthropic0.574—1.0M$5.00 / $25.00proprietary
16SpaceX AI logoGrok 4.7 SpaceX AI0.567—500K$2.00 / $6.00proprietary
17Anthropic logoClaude Opus 4.8 Anthropic0.558—1.0M$5.00 / $25.00proprietary
18OpenAI logoGPT-5.6 Luna OpenAI0.557—1.05M$0.20 / $1.20proprietary
19Meta logoMuse Spark Meta0.541—1.0M—proprietary
20OpenAI logoGPT-5.6 Sol (August) OpenAI0.540—1.05M$4.00 / $20.00proprietary
21Anthropic logoClaude Opus 4.7 Anthropic0.519—1M$5.00 / $25.00proprietary
22OpenAI logoGPT-5.5 OpenAI0.518—1.05M$5.00 / $30.00proprietary
23SpaceX AI logoGrok 4.6 SpaceX AI0.485—500K$2.00 / $6.00proprietary
24OpenAI logoGPT-5.4 OpenAI0.481—1.05M$2.50 / $15.00proprietary
25OpenAI logoGPT-5 OpenAI0.462—400K$1.25 / $10.00proprietary
26OpenAI logoGPT-5.2 OpenAI0.459—400K$1.75 / $14.00proprietary
27Anthropic logoClaude Sonnet 4.6 Anthropic0.442—1M$3.00 / $15.00proprietary
28OpenAI logoGPT-5.6 Luna (August) OpenAI0.441—1.05M$0.20 / $1.20proprietary
29OpenAI logoGPT-5.1 OpenAI0.396—400K$1.25 / $10.00proprietary
30OpenAI logoGPT-5.5 Instant OpenAI0.384—400K$5.00 / $30.00proprietary
31Microsoft logoMAI-Thinking-1 Microsoft0.3501.0T256K$2.00 / $8.00proprietary

Each score is copied as its source document prints it, on the 525-task set. The line under each model name says which document the score was read from and links to its entry on the sources page. List API prices per 1M tokens as of September 30, 2026. GPT-5.6 models charge higher rates above 272K input tokens. MAI-Thinking-1 is in public preview on Microsoft Foundry without final list pricing. Each model name links to its full result.

31
models ranked
5
labs represented
525
physician-authored tasks
15,079
clinician conversations screened
190
contributing physicians
50
countries represented

What the benchmark measures

HealthBench Professional evaluates language models on the work clinicians actually bring to a model: care consults such as differential diagnosis and management questions, clinical writing and documentation, and medical research. The 525 tasks were selected by physicians from 15,079 real clinician conversations, difficult cases were deliberately overweighted, and roughly a third of the set is adversarial. 190 physicians across 50 countries and 26 specialties wrote and adjudicated the rubrics. The full construction is described on the benchmark page.

How to read the scores

Scores are earned rubric points divided by possible points, clipped to 0 to 1, with a length adjustment that keeps verbosity from buying credit. Absolute numbers run low by design: the field's best result is 0.703, and physician-written responses score 0.437 on the same rubrics. Scores depend on the grader version, the reasoning effort, and the deployment setting a document reports, so 2 numbers taken from different documents can differ for reasons other than model quality. The scores here are compiled from published sources, not produced by Arcophos; where each one comes from and how it is read is on the methodology page.

Head to head

Score, price, and context differences for the pairings people actually weigh, with the full score-difference matrix on the compare page.

Data

The full results are published as JSON and CSV at stable URLs, licensed CC BY 4.0. Citation format and version history are on the data page.