GPT-5.6 Luna vs GPT-5.5 Instant on HealthBench Professional
updated August 16, 2026
GPT-5.6 Luna scores 0.557 to GPT-5.5 Instant's 0.384, a gap of 0.173 on the 525 physician-graded tasks of HealthBench Professional. The table below puts the scores next to what each model costs to actually run.
Side by side
| score | 0.557 | 0.384 |
|---|---|---|
| rank | 7 of 9 | 8 of 9 |
| context window | 1.1M | 400K |
| price per 1M tokens, in / out | $0.20 / $1.20 | $5.00 / $30.00 |
| 1,000 consult exchanges | $1.24 | $31.00 |
| released | 2026-07-09 | 2026-05-05 |
| license | proprietary | proprietary |
Consult exchange: 2,000 input and 700 output tokens, priced at list rates as of August 16, 2026. GPT-5.6 models charge higher rates above 272K input tokens. MAI-Thinking-1 is in public preview on Microsoft Foundry without final list pricing.
Reading this pairing
The most lopsided pairing on the board: Luna scores 0.173 higher at a twenty-fifth of the input price. GPT-5.5 Instant predates the 5.6 family and its July 2026 repricing, and this table shows what that generation gap means for healthcare work. Luna is also the cheapest model within 0.11 of the leader, which makes it the default answer for high-volume clinical text at low cost.
Which scores higher on HealthBench Professional, GPT-5.6 Luna or GPT-5.5 Instant?
GPT-5.6 Luna scores higher: 0.557 against GPT-5.5 Instant's 0.384, a difference of 0.173 on the 525-task set, as of August 16, 2026.
Which is cheaper to run, GPT-5.6 Luna or GPT-5.5 Instant?
GPT-5.6 Luna. 1,000 typical consult exchanges (2,000 input and 700 output tokens each) cost $1.24 against $31.00 at list rates.
Full results for both models: GPT-5.6 Luna and GPT-5.5 Instant. The complete score-difference matrix is on the compare page.