Claude Fable 5.1 on HealthBench Professional
rank 3 of 22 · updated September 8, 2026
Claude Fable 5.1 scores 0.621 on HealthBench Professional, rank 3 of 22 models on the board. Anthropic's September 2026 refresh of its top tier; its published score was measured with safety classifiers active and a fallback to Claude Opus 5. HealthBench Professional scores models on 525 tasks drawn from real clinician conversations, graded criterion by criterion against physician-written rubrics on a 0 to 1 scale.
Result and API facts
| rank | 3 of 22 |
|---|---|
| score | 0.621 |
| lab | Anthropic |
| context window | 1.0M tokens |
| API price per 1M tokens | $10.00 in / $50.00 out |
| license | proprietary |
| source | Claude Fable 5.1 and Claude Mythos 5.1 System Card (system card) |
| released | 2026-09-01 |
Position in the field
The gap to the leader, Claude Fable 5 at 0.660, is 0.039. Directly above sits GPT-6 Astra at 0.634. Directly below sits GPT-5.6 Sol at 0.605. The scores on this page are compiled from published documents rather than from one controlled run, so a small gap between 2 models can reflect a difference in grader version, reasoning effort, or deployment setting as well as a difference in capability.
Source of this score
Read from Claude Fable 5.1 and Claude Mythos 5.1 System Card (system card, Anthropic, 2026-09-01). Vendor-reported. Confidence: verified. Configuration: length-adjusted (method published in the HealthBench Professional paper); Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over five trials, no tools or customized system prompt; Fable 5.1 run with safety classifiers active and a refusal-fallback to Claude Opus 5 (raw 74.2%)..
On HealthBench Professional, Claude Fable 5.1 achieved a raw score of 74.2%, ahead of Claude Opus 5 at 73.4%, Fable 5 at 68.9%, and Claude Sonnet 5 at 62.4%. After length adjustment, which penalizes verbose model responses, Fable 5.1 achieved a score of 62.1%.
p. 199, sec. 8.17.2 (same figure printed as 62.1% in Table 8.1.A, p. 167) · full entry on the sources page
What does Claude Fable 5.1 score on HealthBench Professional?
Claude Fable 5.1 scores 0.621 on HealthBench Professional, which places it at rank 3 of 22 models on the board as of September 8, 2026. The number was read from Claude Fable 5.1 and Claude Mythos 5.1 System Card, listed on the sources page.
How much does Claude Fable 5.1 cost to run?
Claude Fable 5.1 is priced at $10.00 per million input tokens and $50.00 per million output tokens through Anthropic's API.
Head to head
Pairings with a dedicated comparison page are linked; every other difference is in the score-difference matrix.
- Claude Fable 5.1 vs Claude Fable 50.621 vs 0.660 · Claude Fable 5 by 0.039
- Claude Fable 5.1 vs GPT-6 Astra0.621 vs 0.634 · GPT-6 Astra by 0.013
- Claude Fable 5.1 vs GPT-5.6 Sol0.621 vs 0.605 · Claude Fable 5.1 by 0.016
- Claude Fable 5.1 vs Claude Opus 50.621 vs 0.598 · Claude Fable 5.1 by 0.023
- Claude Fable 5.1 vs Muse Spark 1.10.621 vs 0.593 · Claude Fable 5.1 by 0.028
- Claude Fable 5.1 vs Claude Sonnet 50.621 vs 0.578 · Claude Fable 5.1 by 0.043
- Claude Fable 5.1 vs GPT-5.6 Terra0.621 vs 0.577 · Claude Fable 5.1 by 0.044
- Claude Fable 5.1 vs Claude Opus 4.80.621 vs 0.558 · Claude Fable 5.1 by 0.063
- Claude Fable 5.1 vs GPT-5.6 Luna0.621 vs 0.557 · Claude Fable 5.1 by 0.064
- Claude Fable 5.1 vs Muse Spark0.621 vs 0.541 · Claude Fable 5.1 by 0.080
- Claude Fable 5.1 vs GPT-5.6 Sol (August)0.621 vs 0.540 · Claude Fable 5.1 by 0.081
- Claude Fable 5.1 vs Claude Opus 4.70.621 vs 0.519 · Claude Fable 5.1 by 0.102
- Claude Fable 5.1 vs GPT-5.50.621 vs 0.518 · Claude Fable 5.1 by 0.103
- Claude Fable 5.1 vs GPT-5.40.621 vs 0.481 · Claude Fable 5.1 by 0.140
- Claude Fable 5.1 vs GPT-50.621 vs 0.462 · Claude Fable 5.1 by 0.159
- Claude Fable 5.1 vs GPT-5.20.621 vs 0.459 · Claude Fable 5.1 by 0.162
- Claude Fable 5.1 vs Claude Sonnet 4.60.621 vs 0.442 · Claude Fable 5.1 by 0.179
- Claude Fable 5.1 vs GPT-5.6 Luna (August)0.621 vs 0.441 · Claude Fable 5.1 by 0.180
- Claude Fable 5.1 vs GPT-5.10.621 vs 0.396 · Claude Fable 5.1 by 0.225
- Claude Fable 5.1 vs GPT-5.5 Instant0.621 vs 0.384 · Claude Fable 5.1 by 0.237
- Claude Fable 5.1 vs MAI-Thinking-10.621 vs 0.350 · Claude Fable 5.1 by 0.271
Where the scores come from and how they are read is on the methodology page. The full ranking is on the leaderboard.