HealthBench Professional

OpenAI logoGPT-5.4 on HealthBench Professional

rank 15 of 22 · updated September 8, 2026

GPT-5.4 scores 0.481 on HealthBench Professional, rank 15 of 22 models on the board. OpenAI's March 2026 model with a 1.05M context window, priced at $2.50 per million input tokens. HealthBench Professional scores models on 525 tasks drawn from real clinician conversations, graded criterion by criterion against physician-written rubrics on a 0 to 1 scale.

Result and API facts

rank15 of 22
score0.481
labOpenAI
context window1.05M tokens
API price per 1M tokens$2.50 in / $15.00 out
licenseproprietary
sourceGPT-5.6 System Card (system card)
released2026-03-05

Position in the field

The gap to the leader, Claude Fable 5 at 0.660, is 0.179. Directly above sits GPT-5.5 at 0.518. Directly below sits GPT-5 at 0.462. The scores on this page are compiled from published documents rather than from one controlled run, so a small gap between 2 models can reflect a difference in grader version, reasoning effort, or deployment setting as well as a difference in capability.

Source of this score

Read from GPT-5.6 System Card (system card, OpenAI, 2026-07-09). Vendor-reported. Confidence: verified. Configuration: length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (51.9 unadjusted, 3308 chars).

HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)

Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.4 · 1 corroborating document · full entry on the sources page

What does GPT-5.4 score on HealthBench Professional?

GPT-5.4 scores 0.481 on HealthBench Professional, which places it at rank 15 of 22 models on the board as of September 8, 2026. The number was read from GPT-5.6 System Card, listed on the sources page.

How much does GPT-5.4 cost to run?

GPT-5.4 is priced at $2.50 per million input tokens and $15.00 per million output tokens through OpenAI's API.

Head to head

Pairings with a dedicated comparison page are linked; every other difference is in the score-difference matrix.

Where the scores come from and how they are read is on the methodology page. The full ranking is on the leaderboard.