# HealthBench Professional Leaderboard, 2026-08-16 Claude Fable 5 leads the HealthBench Professional leaderboard at 0.660, ahead of GPT-5.6 Sol (0.605) and Claude Opus 5 (0.598). 9 models evaluated. Source: https://healthbenchprofessional.com ## Ranking | rank | model | lab | score | context | $ in / 1M | $ out / 1M | license | |---|---|---|---|---|---|---|---| | 1 | Claude Fable 5 | Anthropic | 0.660 | 1.0M | $10.00 | $50.00 | proprietary | | 2 | GPT-5.6 Sol | OpenAI | 0.605 | 1.1M | $5.00 | $30.00 | proprietary | | 3 | Claude Opus 5 | Anthropic | 0.598 | 1.0M | $5.00 | $25.00 | proprietary | | 4 | Claude Sonnet 5 | Anthropic | 0.578 | 1.0M | $2.00 | $10.00 | proprietary | | 5 | GPT-5.6 Terra | OpenAI | 0.577 | 1.1M | $2.00 | $12.00 | proprietary | | 6 | Claude Opus 4.8 | Anthropic | 0.558 | 1.0M | $5.00 | $25.00 | proprietary | | 7 | GPT-5.6 Luna | OpenAI | 0.557 | 1.1M | $0.20 | $1.20 | proprietary | | 8 | GPT-5.5 Instant | OpenAI | 0.384 | 400K | $5.00 | $30.00 | proprietary | | 9 | MAI-Thinking-1 | Microsoft | 0.350 | - | - | - | proprietary | Prices are vendor list rates per 1M tokens as of 2026-08-16. GPT-5.6 models charge higher rates above 272K input tokens. MAI-Thinking-1 is in public preview on Microsoft Foundry without final list pricing. ## About the benchmark HealthBench Professional is a benchmark published by OpenAI in April 2026 (arXiv 2604.27470). It contains 525 physician-authored tasks selected from 15,079 real clinician conversations, covering care consults, clinical writing and documentation, and medical research. 190 physicians across 50 countries and 26 specialties wrote and adjudicated the rubrics; roughly a third of the tasks are adversarial. A model grader (GPT-5.4, low reasoning effort) judges each physician-written rubric criterion as met or not met. The score is earned points divided by possible points, clipped to 0 to 1, with a length adjustment so verbose answers cannot buy credit. Physician-written responses score 0.437 on the same rubrics. ## How these numbers were produced Each model was run through its vendor's public API under default sampling, one response per task, no tools and no scaffolding, graded with the benchmark's default grader configuration. Scores are comparable within this table; they are not comparable against numbers graded under other configurations, including OpenAI's in-product results. ## Models - Claude Fable 5 (Anthropic): 0.660, rank 1 of 9. Anthropic's top-tier model, positioned above the Opus line for the hardest reasoning work. Released 2026-06-09. https://healthbenchprofessional.com/models/claude-fable-5 - GPT-5.6 Sol (OpenAI): 0.605, rank 2 of 9. The flagship tier of OpenAI's GPT-5.6 family, aimed at its hardest reasoning and agentic workloads. Released 2026-07-09. https://healthbenchprofessional.com/models/gpt-5.6-sol - Claude Opus 5 (Anthropic): 0.598, rank 3 of 9. Anthropic's frontier workhorse, priced at half of Claude Fable 5. Released 2026-07-24. https://healthbenchprofessional.com/models/claude-opus-5 - Claude Sonnet 5 (Anthropic): 0.578, rank 4 of 9. Anthropic's mid-tier model, positioned for speed at near-Opus quality. Released 2026-06-30. https://healthbenchprofessional.com/models/claude-sonnet-5 - GPT-5.6 Terra (OpenAI): 0.577, rank 5 of 9. The balanced middle tier of the GPT-5.6 family, between Sol and Luna. Released 2026-07-09. https://healthbenchprofessional.com/models/gpt-5.6-terra - Claude Opus 4.8 (Anthropic): 0.558, rank 6 of 9. The last release of the Opus 4 series, superseded by Claude Opus 5 in July 2026. Released 2026-05-28. https://healthbenchprofessional.com/models/claude-opus-4.8 - GPT-5.6 Luna (OpenAI): 0.557, rank 7 of 9. The small, high-volume tier of the GPT-5.6 family, repriced to $0.20 per million input tokens on July 30, 2026. Released 2026-07-09. https://healthbenchprofessional.com/models/gpt-5.6-luna - GPT-5.5 Instant (OpenAI): 0.384, rank 8 of 9. The fast variant that became ChatGPT's default model in May 2026. Released 2026-05-05. https://healthbenchprofessional.com/models/gpt-5.5-instant - MAI-Thinking-1 (Microsoft): 0.350, rank 9 of 9. Microsoft AI's first in-house frontier reasoning model, a sparse mixture-of-experts design in public preview on Microsoft Foundry. Released 2026-08-12. https://healthbenchprofessional.com/models/mai-thinking-1 ## Reuse Data: CC BY 4.0. Cite as: Arcophos. HealthBench Professional Leaderboard. 2026-08-16 snapshot. https://healthbenchprofessional.com. Machine-readable: https://healthbenchprofessional.com/data/leaderboard.json