HealthBench Professional

Anthropic logoClaude Sonnet 5 on HealthBench Professional

rank 13 of 31 · updated September 30, 2026

Claude Sonnet 5 scores 0.578 on HealthBench Professional, rank 13 of 31 models on the board. Anthropic's mid-tier model, positioned for speed at near-Opus quality. HealthBench Professional scores models on 525 tasks drawn from real clinician conversations, graded criterion by criterion against physician-written rubrics on a 0 to 1 scale.

Result and API facts

rank13 of 31
score0.578
labAnthropic
context window1.0M tokens
API price per 1M tokens$2.00 in / $10.00 out
licenseproprietary
sourceSystem Card: Claude Sonnet 5 (system card)
released2026-06-30

Position in the field

The gap to the leader, GPT-6 Astra (Anthropic run) at 0.703, is 0.125. Directly above sits Muse Spark 1.1 at 0.593. Directly below sits GPT-5.6 Terra at 0.577. The scores on this page are compiled from published documents rather than from one controlled run, so a small gap between 2 models can reflect a difference in grader version, reasoning effort, or deployment setting as well as a difference in capability.

Source of this score

Read from System Card: Claude Sonnet 5 (system card, Anthropic, 2026-06-30). Vendor-reported. Confidence: verified. Configuration: length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 62.4%)..

HealthBench Professional 57.8 44.2 51.8 -

p. 115, Table 8.1.A, row HealthBench Professional, column Claude Sonnet 5; Figure 8.12.2.A p. 139 · 1 corroborating document · full entry on the sources page

What does Claude Sonnet 5 score on HealthBench Professional?

Claude Sonnet 5 scores 0.578 on HealthBench Professional, which places it at rank 13 of 31 models on the board as of September 30, 2026. The number was read from System Card: Claude Sonnet 5, listed on the sources page.

How much does Claude Sonnet 5 cost to run?

Claude Sonnet 5 is priced at $2.00 per million input tokens and $10.00 per million output tokens through Anthropic's API.

Head to head

Pairings with a dedicated comparison page are linked; every other difference is in the score-difference matrix.

Where the scores come from and how they are read is on the methodology page. The full ranking is on the leaderboard.