HealthBench Professional

Microsoft logoMAI-Thinking-1 on HealthBench Professional

rank 31 of 31 · updated September 30, 2026

MAI-Thinking-1 scores 0.350 on HealthBench Professional, rank 31 of 31 models on the board. Microsoft AI's first in-house frontier reasoning model, a sparse mixture-of-experts design in public preview on Microsoft Foundry. HealthBench Professional scores models on 525 tasks drawn from real clinician conversations, graded criterion by criterion against physician-written rubrics on a 0 to 1 scale.

Result and API facts

rank31 of 31
score0.350
labMicrosoft
context window256K tokens
API price per 1M tokens$2.00 in / $8.00 out
licenseproprietary
sourceMAI-Thinking-1: Building a Hill-Climbing Machine (model card)
parameters1.0T
released2026-08-12

Position in the field

The gap to the leader, GPT-6 Astra (Anthropic run) at 0.703, is 0.353. Directly above sits GPT-5.5 Instant at 0.384. The scores on this page are compiled from published documents rather than from one controlled run, so a small gap between 2 models can reflect a difference in grader version, reasoning effort, or deployment setting as well as a difference in capability.

Source of this score

Read from MAI-Thinking-1: Building a Hill-Climbing Machine (model card, Microsoft AI, 2026-08-12). Vendor-reported. Confidence: verified. Configuration: length-adjusted (HealthBench Professional length penalty), standard GPT-5.4 grader and OpenAI rubrics, Microsoft AI run; printed at integer precision..

Model AIR-Bench CyberSec Instruct CyberSec Auto Long Fact Truthful QA HealthBench Prof. MedXpert QA MAI-Thinking-1 88 63 63 98 88 35 43 Sonnet 4.6 88 62 56 98 88 38 49

p. 54, Table 12 'Post-trained model evaluation results on various public benchmarks', Health group, column HealthBench Prof.; protocol Appendix K.6 p. 106: HealthBench Professional introduces a length penalty for the primary metric, to correct for a well-observed correlation between lengthy responses and artificially increased LLM-grader scores. For all reported scores, we use the standard GPT-5.4 grader and rubrics provided by OpenAI. · full entry on the sources page

What does MAI-Thinking-1 score on HealthBench Professional?

MAI-Thinking-1 scores 0.350 on HealthBench Professional, which places it at rank 31 of 31 models on the board as of September 30, 2026. The number was read from MAI-Thinking-1: Building a Hill-Climbing Machine, listed on the sources page.

How much does MAI-Thinking-1 cost to run?

MAI-Thinking-1 is priced at $2.00 per million input tokens and $8.00 per million output tokens through Microsoft's API.

Head to head

Pairings with a dedicated comparison page are linked; every other difference is in the score-difference matrix.

Where the scores come from and how they are read is on the methodology page. The full ranking is on the leaderboard.