Models
9 models evaluated · updated August 16, 2026
Every model on the HealthBench Professional leaderboard, with the API facts that decide whether a score is usable in practice: context window, per-token price, and license. Each model has its own page with the full result and head-to-head comparisons.
Ranking
| # | model | score | size | context | cost in / out per 1M | license | |
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 Anthropic | 0.660 | — | 1.0M | $10.00 / $50.00 | proprietary | |
| 2 | GPT-5.6 Sol OpenAI | 0.605 | — | 1.1M | $5.00 / $30.00 | proprietary | |
| 3 | Claude Opus 5 Anthropic | 0.598 | — | 1.0M | $5.00 / $25.00 | proprietary | |
| 4 | Claude Sonnet 5 Anthropic | 0.578 | — | 1.0M | $2.00 / $10.00 | proprietary | |
| 5 | GPT-5.6 Terra OpenAI | 0.577 | — | 1.1M | $2.00 / $12.00 | proprietary | |
| 6 | Claude Opus 4.8 Anthropic | 0.558 | — | 1.0M | $5.00 / $25.00 | proprietary | |
| 7 | GPT-5.6 Luna OpenAI | 0.557 | — | 1.1M | $0.20 / $1.20 | proprietary | |
| 8 | GPT-5.5 Instant OpenAI | 0.384 | — | 400K | $5.00 / $30.00 | proprietary | |
| 9 | MAI-Thinking-1 Microsoft | 0.350 | 1.0T | — | — | proprietary | |
Model pages
- rank 1 of 9 · score 0.660 · Anthropic
- rank 2 of 9 · score 0.605 · OpenAI
- rank 3 of 9 · score 0.598 · Anthropic
- rank 4 of 9 · score 0.578 · Anthropic
- rank 5 of 9 · score 0.577 · OpenAI
- rank 6 of 9 · score 0.558 · Anthropic
- rank 7 of 9 · score 0.557 · OpenAI
- rank 8 of 9 · score 0.384 · OpenAI
- rank 9 of 9 · score 0.350 · Microsoft