HealthBench Professional

Anthropic logoClaude Sonnet 5.5 on HealthBench Professional

rank 2 of 31 · updated September 30, 2026

Claude Sonnet 5.5 scores 0.692 on HealthBench Professional, rank 2 of 31 models on the board. HealthBench Professional scores models on 525 tasks drawn from real clinician conversations, graded criterion by criterion against physician-written rubrics on a 0 to 1 scale.

Result and API facts

rank2 of 31
score0.692
labAnthropic
context window1M tokens
API price per 1M tokens$2.00 in / $10.00 out
licenseproprietary
sourceClaude Sonnet 5.5 System Card (system card)
released2026-09-28

Position in the field

The gap to the leader, GPT-6 Astra (Anthropic run) at 0.703, is 0.011. Directly below sits Claude Fable 5 at 0.660. The scores on this page are compiled from published documents rather than from one controlled run, so a small gap between 2 models can reflect a difference in grader version, reasoning effort, or deployment setting as well as a difference in capability.

Source of this score

Read from Claude Sonnet 5.5 System Card (system card, Anthropic, 2026-09-28). Vendor-reported. Confidence: verified. Configuration: Anthropic; max effort; Opus 4.8 grader; safety classifiers enabled; length-adjusted; raw 77.1%.

On HealthBench Professional at max effort, length adjustment changes the ranking. After: GPT-6 Astra (70.3%) > Claude Sonnet 5.5 (69.2%) > Claude Opus 5.5 (65.6%) > Claude Fable 5.1 (62.1%).

p. 138, Section 8.15.2 and paragraph above Figure 8.15.B. · full entry on the sources page

What does Claude Sonnet 5.5 score on HealthBench Professional?

Claude Sonnet 5.5 scores 0.692 on HealthBench Professional, which places it at rank 2 of 31 models on the board as of September 30, 2026. The number was read from Claude Sonnet 5.5 System Card, listed on the sources page.

How much does Claude Sonnet 5.5 cost to run?

Claude Sonnet 5.5 is priced at $2.00 per million input tokens and $10.00 per million output tokens through Anthropic's API.

Head to head

Pairings with a dedicated comparison page are linked; every other difference is in the score-difference matrix.

Where the scores come from and how they are read is on the methodology page. The full ranking is on the leaderboard.