GPT-5.1 on HealthBench Professional
rank 20 of 22 · updated September 8, 2026
GPT-5.1 scores 0.396 on HealthBench Professional, rank 20 of 22 models on the board. OpenAI's November 2025 revision of GPT-5, priced like GPT-5 with a 400K context window. HealthBench Professional scores models on 525 tasks drawn from real clinician conversations, graded criterion by criterion against physician-written rubrics on a 0 to 1 scale.
Result and API facts
| rank | 20 of 22 |
|---|---|
| score | 0.396 |
| lab | OpenAI |
| context window | 400K tokens |
| API price per 1M tokens | $1.25 in / $10.00 out |
| license | proprietary |
| source | GPT-5.6 System Card (system card) |
| released | 2025-11-12 |
Position in the field
The gap to the leader, Claude Fable 5 at 0.660, is 0.264. Directly above sits GPT-5.6 Luna (August) at 0.441. Directly below sits GPT-5.5 Instant at 0.384. The scores on this page are compiled from published documents rather than from one controlled run, so a small gap between 2 models can reflect a difference in grader version, reasoning effort, or deployment setting as well as a difference in capability.
Source of this score
Read from GPT-5.6 System Card (system card, OpenAI, 2026-07-09). Vendor-reported. Confidence: verified. Configuration: length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.1 (48.0 unadjusted, 4863 chars).
HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.1 · 1 corroborating document · full entry on the sources page
What does GPT-5.1 score on HealthBench Professional?
GPT-5.1 scores 0.396 on HealthBench Professional, which places it at rank 20 of 22 models on the board as of September 8, 2026. The number was read from GPT-5.6 System Card, listed on the sources page.
How much does GPT-5.1 cost to run?
GPT-5.1 is priced at $1.25 per million input tokens and $10.00 per million output tokens through OpenAI's API.
Head to head
Pairings with a dedicated comparison page are linked; every other difference is in the score-difference matrix.
- GPT-5.1 vs Claude Fable 50.396 vs 0.660 · Claude Fable 5 by 0.264
- GPT-5.1 vs GPT-6 Astra0.396 vs 0.634 · GPT-6 Astra by 0.238
- GPT-5.1 vs Claude Fable 5.10.396 vs 0.621 · Claude Fable 5.1 by 0.225
- GPT-5.1 vs GPT-5.6 Sol0.396 vs 0.605 · GPT-5.6 Sol by 0.209
- GPT-5.1 vs Claude Opus 50.396 vs 0.598 · Claude Opus 5 by 0.202
- GPT-5.1 vs Muse Spark 1.10.396 vs 0.593 · Muse Spark 1.1 by 0.197
- GPT-5.1 vs Claude Sonnet 50.396 vs 0.578 · Claude Sonnet 5 by 0.182
- GPT-5.1 vs GPT-5.6 Terra0.396 vs 0.577 · GPT-5.6 Terra by 0.181
- GPT-5.1 vs Claude Opus 4.80.396 vs 0.558 · Claude Opus 4.8 by 0.162
- GPT-5.1 vs GPT-5.6 Luna0.396 vs 0.557 · GPT-5.6 Luna by 0.161
- GPT-5.1 vs Muse Spark0.396 vs 0.541 · Muse Spark by 0.145
- GPT-5.1 vs GPT-5.6 Sol (August)0.396 vs 0.540 · GPT-5.6 Sol (August) by 0.144
- GPT-5.1 vs Claude Opus 4.70.396 vs 0.519 · Claude Opus 4.7 by 0.123
- GPT-5.1 vs GPT-5.50.396 vs 0.518 · GPT-5.5 by 0.122
- GPT-5.1 vs GPT-5.40.396 vs 0.481 · GPT-5.4 by 0.085
- GPT-5.1 vs GPT-50.396 vs 0.462 · GPT-5 by 0.066
- GPT-5.1 vs GPT-5.20.396 vs 0.459 · GPT-5.2 by 0.063
- GPT-5.1 vs Claude Sonnet 4.60.396 vs 0.442 · Claude Sonnet 4.6 by 0.046
- GPT-5.1 vs GPT-5.6 Luna (August)0.396 vs 0.441 · GPT-5.6 Luna (August) by 0.045
- GPT-5.1 vs GPT-5.5 Instant0.396 vs 0.384 · GPT-5.1 by 0.012
- GPT-5.1 vs MAI-Thinking-10.396 vs 0.350 · GPT-5.1 by 0.046
Where the scores come from and how they are read is on the methodology page. The full ranking is on the leaderboard.