MAI-Thinking-1 on HealthBench Professional
rank 9 of 9 · updated August 16, 2026
MAI-Thinking-1 scores 0.350 on HealthBench Professional, rank 9 of 9 evaluated models. Microsoft AI's first in-house frontier reasoning model, a sparse mixture-of-experts design in public preview on Microsoft Foundry. HealthBench Professional scores models on 525 tasks drawn from real clinician conversations, graded criterion by criterion against physician-written rubrics on a 0 to 1 scale.
Result and API facts
| rank | 9 of 9 |
|---|---|
| score | 0.350 |
| lab | Microsoft |
| context window | not published |
| API price per 1M tokens | no published list pricing |
| license | proprietary |
| parameters | 1.0T |
| released | 2026-08-12 |
Position in the field
The gap to the leader, Claude Fable 5 at 0.660, is 0.310. Directly above sits GPT-5.5 Instant at 0.384. Scores on this page come from the same evaluation run, so differences between models are differences on identical tasks, not across configurations.
What does MAI-Thinking-1 score on HealthBench Professional?
MAI-Thinking-1 scores 0.350 on HealthBench Professional, which places it at rank 9 of 9 evaluated models as of August 16, 2026.
How much does MAI-Thinking-1 cost to run?
Microsoft has not published API pricing for MAI-Thinking-1.
Head to head
Pairings with a dedicated comparison page are linked; every other difference is in the score-difference matrix.
- MAI-Thinking-1 vs Claude Fable 50.350 vs 0.660 · Claude Fable 5 by 0.310
- MAI-Thinking-1 vs GPT-5.6 Sol0.350 vs 0.605 · GPT-5.6 Sol by 0.255
- MAI-Thinking-1 vs Claude Opus 50.350 vs 0.598 · Claude Opus 5 by 0.248
- MAI-Thinking-1 vs Claude Sonnet 50.350 vs 0.578 · Claude Sonnet 5 by 0.228
- MAI-Thinking-1 vs GPT-5.6 Terra0.350 vs 0.577 · GPT-5.6 Terra by 0.227
- MAI-Thinking-1 vs Claude Opus 4.80.350 vs 0.558 · Claude Opus 4.8 by 0.208
- MAI-Thinking-1 vs GPT-5.6 Luna0.350 vs 0.557 · GPT-5.6 Luna by 0.207
- MAI-Thinking-1 vs GPT-5.5 Instant0.350 vs 0.384 · GPT-5.5 Instant by 0.034
How tasks are selected and graded is on the methodology page. The full ranking is on the leaderboard.