MAI-Thinking-1 on HealthBench Professional
rank 31 of 31 · updated September 30, 2026
MAI-Thinking-1 scores 0.350 on HealthBench Professional, rank 31 of 31 models on the board. Microsoft AI's first in-house frontier reasoning model, a sparse mixture-of-experts design in public preview on Microsoft Foundry. HealthBench Professional scores models on 525 tasks drawn from real clinician conversations, graded criterion by criterion against physician-written rubrics on a 0 to 1 scale.
Result and API facts
| rank | 31 of 31 |
|---|---|
| score | 0.350 |
| lab | Microsoft |
| context window | 256K tokens |
| API price per 1M tokens | $2.00 in / $8.00 out |
| license | proprietary |
| source | MAI-Thinking-1: Building a Hill-Climbing Machine (model card) |
| parameters | 1.0T |
| released | 2026-08-12 |
Position in the field
The gap to the leader, GPT-6 Astra (Anthropic run) at 0.703, is 0.353. Directly above sits GPT-5.5 Instant at 0.384. The scores on this page are compiled from published documents rather than from one controlled run, so a small gap between 2 models can reflect a difference in grader version, reasoning effort, or deployment setting as well as a difference in capability.
Source of this score
Read from MAI-Thinking-1: Building a Hill-Climbing Machine (model card, Microsoft AI, 2026-08-12). Vendor-reported. Confidence: verified. Configuration: length-adjusted (HealthBench Professional length penalty), standard GPT-5.4 grader and OpenAI rubrics, Microsoft AI run; printed at integer precision..
Model AIR-Bench CyberSec Instruct CyberSec Auto Long Fact Truthful QA HealthBench Prof. MedXpert QA MAI-Thinking-1 88 63 63 98 88 35 43 Sonnet 4.6 88 62 56 98 88 38 49
p. 54, Table 12 'Post-trained model evaluation results on various public benchmarks', Health group, column HealthBench Prof.; protocol Appendix K.6 p. 106: HealthBench Professional introduces a length penalty for the primary metric, to correct for a well-observed correlation between lengthy responses and artificially increased LLM-grader scores. For all reported scores, we use the standard GPT-5.4 grader and rubrics provided by OpenAI. · full entry on the sources page
What does MAI-Thinking-1 score on HealthBench Professional?
MAI-Thinking-1 scores 0.350 on HealthBench Professional, which places it at rank 31 of 31 models on the board as of September 30, 2026. The number was read from MAI-Thinking-1: Building a Hill-Climbing Machine, listed on the sources page.
How much does MAI-Thinking-1 cost to run?
MAI-Thinking-1 is priced at $2.00 per million input tokens and $8.00 per million output tokens through Microsoft's API.
Head to head
Pairings with a dedicated comparison page are linked; every other difference is in the score-difference matrix.
- MAI-Thinking-1 vs GPT-6 Astra (Anthropic run)0.350 vs 0.703 · GPT-6 Astra (Anthropic run) by 0.353
- MAI-Thinking-1 vs Claude Sonnet 5.50.350 vs 0.692 · Claude Sonnet 5.5 by 0.342
- MAI-Thinking-1 vs Claude Fable 50.350 vs 0.660 · Claude Fable 5 by 0.310
- MAI-Thinking-1 vs Claude Opus 5.50.350 vs 0.656 · Claude Opus 5.5 by 0.306
- MAI-Thinking-1 vs GPT-6 Astra0.350 vs 0.647 · GPT-6 Astra by 0.297
- MAI-Thinking-1 vs Claude Fable 5 (September card)0.350 vs 0.633 · Claude Fable 5 (September card) by 0.283
- MAI-Thinking-1 vs Claude Fable 5.10.350 vs 0.621 · Claude Fable 5.1 by 0.271
- MAI-Thinking-1 vs GPT-6 Luna0.350 vs 0.608 · GPT-6 Luna by 0.258
- MAI-Thinking-1 vs GPT-6 Sol0.350 vs 0.608 · GPT-6 Sol by 0.258
- MAI-Thinking-1 vs GPT-5.6 Sol0.350 vs 0.605 · GPT-5.6 Sol by 0.255
- MAI-Thinking-1 vs Claude Opus 50.350 vs 0.598 · Claude Opus 5 by 0.248
- MAI-Thinking-1 vs Muse Spark 1.10.350 vs 0.593 · Muse Spark 1.1 by 0.243
- MAI-Thinking-1 vs Claude Sonnet 50.350 vs 0.578 · Claude Sonnet 5 by 0.228
- MAI-Thinking-1 vs GPT-5.6 Terra0.350 vs 0.577 · GPT-5.6 Terra by 0.227
- MAI-Thinking-1 vs Claude Opus 4.8 (Opus 4.8 grader)0.350 vs 0.574 · Claude Opus 4.8 (Opus 4.8 grader) by 0.224
- MAI-Thinking-1 vs Grok 4.70.350 vs 0.567 · Grok 4.7 by 0.217
- MAI-Thinking-1 vs Claude Opus 4.80.350 vs 0.558 · Claude Opus 4.8 by 0.208
- MAI-Thinking-1 vs GPT-5.6 Luna0.350 vs 0.557 · GPT-5.6 Luna by 0.207
- MAI-Thinking-1 vs Muse Spark0.350 vs 0.541 · Muse Spark by 0.191
- MAI-Thinking-1 vs GPT-5.6 Sol (August)0.350 vs 0.540 · GPT-5.6 Sol (August) by 0.190
- MAI-Thinking-1 vs Claude Opus 4.70.350 vs 0.519 · Claude Opus 4.7 by 0.169
- MAI-Thinking-1 vs GPT-5.50.350 vs 0.518 · GPT-5.5 by 0.168
- MAI-Thinking-1 vs Grok 4.60.350 vs 0.485 · Grok 4.6 by 0.135
- MAI-Thinking-1 vs GPT-5.40.350 vs 0.481 · GPT-5.4 by 0.131
- MAI-Thinking-1 vs GPT-50.350 vs 0.462 · GPT-5 by 0.112
- MAI-Thinking-1 vs GPT-5.20.350 vs 0.459 · GPT-5.2 by 0.109
- MAI-Thinking-1 vs Claude Sonnet 4.60.350 vs 0.442 · Claude Sonnet 4.6 by 0.092
- MAI-Thinking-1 vs GPT-5.6 Luna (August)0.350 vs 0.441 · GPT-5.6 Luna (August) by 0.091
- MAI-Thinking-1 vs GPT-5.10.350 vs 0.396 · GPT-5.1 by 0.046
- MAI-Thinking-1 vs GPT-5.5 Instant0.350 vs 0.384 · GPT-5.5 Instant by 0.034
Where the scores come from and how they are read is on the methodology page. The full ranking is on the leaderboard.