HealthBench Professional

The best AI model for healthcare, September 2026

based on the HealthBench Professional leaderboard, updated September 30, 2026

Measured on 525 physician-graded clinical tasks, Claude Fable 5 is the strongest model for professional healthcare work right now, at 0.660. It is also the most expensive, and the right pick changes with the constraint you are actually under, so the table below answers the question four ways. The numbers come from the HealthBench Professional leaderboard, which this site compiles from published vendor documents and leaderboards; each row links to its source.

The short version

Highest score, price no objectClaude Fable 50.660, ahead of the field by 0.055. $10.00 / $50.00 per 1M tokens.
The $5 tierGPT-5.6 Sol or Claude Opus 50.605 vs 0.598, a 0.007 gap that reads as a tie. Opus 5's output tokens are $-5 per 1M cheaper.
The $2 tierClaude Sonnet 5 or GPT-5.6 Terra0.578 vs 0.577, 0.001 apart. Sonnet 5's output tokens are $2 per 1M cheaper.
High volume on a budgetGPT-5.6 Luna0.557 at $0.20 / $1.20 per 1M tokens, within 0.103 of the leader at a fraction of every other price on the board.

Why these picks

The top of the board splits into price tiers, and inside two of them the scores are effectively tied: 0.007 separates the $5-tier pair and 0.001 separates the $2-tier pair, both smaller than the differences that grader version or reasoning effort alone can introduce between 2 published numbers. That makes the honest recommendation at those tiers a toss-up settled by output pricing and your existing stack, not by the benchmark. The two decisions the benchmark does settle: Claude Fable 5's lead over everything else is real, and GPT-5.6 Luna delivers most of the field's capability for about a twenty-fifth of the $5-tier input price.

Where are the open-weights models?

Every model currently on the board ships behind a proprietary API; no open-weights model has a published HealthBench Professional result on this board yet. When one is added, this page will carry a pick for deployments that require local weights, which is a common constraint for health systems handling records on-premises.

What this page is not

A benchmark score measures rubric adherence on hard, written tasks. It does not measure bedside judgment, integration effort, latency, or whether a vendor will sign the agreements your compliance team needs, and none of it is medical advice. Read the methodology before treating any single number as a procurement decision, and read the benchmark page to understand what the tasks actually contain.