HealthBench Professional

Claude Sonnet 5 vs GPT-5.6 Terra on HealthBench Professional

updated August 16, 2026

Claude Sonnet 5 scores 0.578 to GPT-5.6 Terra's 0.577, a gap of 0.001 on the 525 physician-graded tasks of HealthBench Professional. The table below puts the scores next to what each model costs to actually run.

Side by side

Anthropic logoClaude Sonnet 5OpenAI logoGPT-5.6 Terra
score0.5780.577
rank4 of 95 of 9
context window1.0M1.1M
price per 1M tokens, in / out$2.00 / $10.00$2.00 / $12.00
1,000 consult exchanges$11.00$12.40
released2026-06-302026-07-09
licenseproprietaryproprietary

Consult exchange: 2,000 input and 700 output tokens, priced at list rates as of August 16, 2026. GPT-5.6 models charge higher rates above 272K input tokens. MAI-Thinking-1 is in public preview on Microsoft Foundry without final list pricing.

Reading this pairing

The $2 tier decides most real deployments, and it is a dead heat: 0.001 separates Sonnet 5 and Terra, tied for the closest pairing on the board. Sonnet's output tokens are $2 per million cheaper, which at documentation-workload volumes is the only difference that compounds. Treat the scores as a tie and choose on price and integration.

Which scores higher on HealthBench Professional, Claude Sonnet 5 or GPT-5.6 Terra?

Claude Sonnet 5 scores higher: 0.578 against GPT-5.6 Terra's 0.577, a difference of 0.001 on the 525-task set, as of August 16, 2026.

Which is cheaper to run, Claude Sonnet 5 or GPT-5.6 Terra?

Claude Sonnet 5. 1,000 typical consult exchanges (2,000 input and 700 output tokens each) cost $11.00 against $12.40 at list rates.

Full results for both models: Claude Sonnet 5 and GPT-5.6 Terra. The complete score-difference matrix is on the compare page.