HealthBench Professional

Anthropic logoClaude Opus 4.8 (Opus 4.8 grader) on HealthBench Professional

rank 15 of 31 · updated September 30, 2026

Claude Opus 4.8 (Opus 4.8 grader) scores 0.574 on HealthBench Professional, rank 15 of 31 models on the board. The last release of the Opus 4 series, superseded by Claude Opus 5 in July 2026. HealthBench Professional scores models on 525 tasks drawn from real clinician conversations, graded criterion by criterion against physician-written rubrics on a 0 to 1 scale.

Result and API facts

rank15 of 31
score0.574
labAnthropic
context window1.0M tokens
API price per 1M tokens$5.00 in / $25.00 out
licenseproprietary
sourceSystem Card: Claude Sonnet 5 (system card)
released2026-05-28

Position in the field

The gap to the leader, GPT-6 Astra (Anthropic run) at 0.703, is 0.129. Directly above sits GPT-5.6 Terra at 0.577. Directly below sits Grok 4.7 at 0.567. The scores on this page are compiled from published documents rather than from one controlled run, so a small gap between 2 models can reflect a difference in grader version, reasoning effort, or deployment setting as well as a difference in capability.

Source of this score

Read from System Card: Claude Sonnet 5 (system card, Anthropic, 2026-06-30). Vendor-reported. Confidence: verified. Configuration: Anthropic June evaluation; length-adjusted; adaptive max effort; Opus 4.8 grader; five trials; no tools or custom system prompt.

HealthBench Professional | Length-adjusted score (%) | Claude Opus 4.8 | 57.4%

p. 139, Figure 8.12.2.A, Claude Opus 4.8 bar (visually read printed labels). · full entry on the sources page

What does Claude Opus 4.8 (Opus 4.8 grader) score on HealthBench Professional?

Claude Opus 4.8 (Opus 4.8 grader) scores 0.574 on HealthBench Professional, which places it at rank 15 of 31 models on the board as of September 30, 2026. The number was read from System Card: Claude Sonnet 5, listed on the sources page.

How much does Claude Opus 4.8 (Opus 4.8 grader) cost to run?

Claude Opus 4.8 (Opus 4.8 grader) is priced at $5.00 per million input tokens and $25.00 per million output tokens through Anthropic's API.

Head to head

Pairings with a dedicated comparison page are linked; every other difference is in the score-difference matrix.

Where the scores come from and how they are read is on the methodology page. The full ranking is on the leaderboard.