HealthBench Professional

GPT-5.6 Luna vs GPT-5.5 Instant on HealthBench Professional

updated September 30, 2026

GPT-5.6 Luna scores 0.557 to GPT-5.5 Instant's 0.384, a gap of 0.173 on the 525 physician-graded tasks of HealthBench Professional. The table below puts the scores next to what each model costs to actually run.

Side by side

OpenAI logoGPT-5.6 LunaOpenAI logoGPT-5.5 Instant
score0.5570.384
rank18 of 3130 of 31
context window1.05M400K
price per 1M tokens, in / out$0.20 / $1.20$5.00 / $30.00
1,000 consult exchanges$1.24$31.00
released2026-07-092026-05-05
licenseproprietaryproprietary

Consult exchange: 2,000 input and 700 output tokens, priced at list rates as of September 30, 2026. GPT-5.6 models charge higher rates above 272K input tokens. MAI-Thinking-1 is in public preview on Microsoft Foundry without final list pricing.

Reading this pairing

The most lopsided pairing on the board: Luna scores 0.173 higher at a twenty-fifth of the input price. GPT-5.5 Instant predates the 5.6 family and its July 2026 repricing, and this table shows what that generation gap means for healthcare work. Luna is also the cheapest model within 0.11 of the leader, which makes it the default answer for high-volume clinical text at low cost.

Which scores higher on HealthBench Professional, GPT-5.6 Luna or GPT-5.5 Instant?

GPT-5.6 Luna scores higher: 0.557 against GPT-5.5 Instant's 0.384, a difference of 0.173 on the 525-task set, as of September 30, 2026.

Which is cheaper to run, GPT-5.6 Luna or GPT-5.5 Instant?

GPT-5.6 Luna. 1,000 typical consult exchanges (2,000 input and 700 output tokens each) cost $1.24 against $31.00 at list rates.

Full results for both models: GPT-5.6 Luna and GPT-5.5 Instant. The complete score-difference matrix is on the compare page.