HealthBench Professional

Methodology

The scores on this site are compiled from published documents. Arcophos does not currently run HealthBench Professional itself. Each number was read from a vendor system card, an official update report, a model card or launch announcement, or a named public leaderboard, and every row links to the document it was taken from on the sources page. Rows for which no document has been located yet are marked source pending on the board and listed separately there. What the benchmark measures and how it was constructed is on the benchmark page; this page covers where the numbers come from and how they are read.

Where the numbers come from

The board currently draws on 17 documents for 31 rows. Every row has a located document. First-party documents take priority: the vendor's own system card for a vendor-reported number, the benchmark maintainers' leaderboard for a board number. Pages that repeat those numbers are treated as corroboration only.

How numbers are read

Task setThe released 525-task HealthBench Professional set. Task construction is OpenAI's and is described on the benchmark page. A document that reports a subset or a different task count is not used for the main score.
DocumentsOpenAI system cards and update reports, Anthropic system cards, other vendors' model cards and launch announcements, and named public leaderboards. A mirror or press article can corroborate a number but is not its source.
CopyingThe score is copied exactly as the document prints it, with the same number of decimals. Each row stores the verbatim sentence or table cell the number was read from and a locator such as a page or table reference.
PrecedenceWhen 2 documents disagree, the row keeps the number from the higher-authority document: system or model card first, then an official leaderboard, then a paper, then a launch post, then an independent leaderboard. The disagreement is kept on record and shown on the sources page.
ConfidenceVerified: the number was read from a first-party document or an official leaderboard. Partial: the number was read from a lower-authority document such as an independent leaderboard or a paper by a third party. Unverified: no document has been located yet. Disputed: located documents disagree and the difference is unresolved. Unverified and disputed rows are marked source pending on the board.
Pricing and specsContext windows and per-token prices come from vendor documentation and list price sheets, checked on the update date. Prices shown are base rates as of September 30, 2026.

Why documents disagree

A compiled board puts numbers from different documents in one column, and those documents were not produced under one configuration. 4 settings account for most of the differences between them.

Length adjustmentThe benchmark applies a length penalty so long answers cannot buy credit. A document that reports the raw rubric score without the adjustment prints a different number from one that applies it, for the same responses.
Reasoning effortVendors report scores at a chosen reasoning effort, and the same model at low and high effort does not score the same. A document does not always say which setting it used.
Deployment settingOpenAI's paper reports GPT-5.4 inside ChatGPT for Clinicians separately from base GPT-5.4. The product setting carries scaffolding the base model does not have, and the 2 numbers are not interchangeable.
Grader versionRubric scores are model-graded. The reference grader is GPT-5.4, low reasoning effort, but a document graded with an earlier or later grader will place the same responses at a different score.

For that reason a gap of a few thousandths between 2 rows read from different documents is not evidence that one model is better. The sources page shows the excerpt behind each number so the reader can check what each document actually measured.

Limitations

A compiled board is not a controlled comparison. Vendors choose which numbers to publish and under which settings, and a model with no published result is absent rather than scored low. The grader is itself a model, and grader bias is a known open problem for rubric benchmarks. Per-use-case subscores (care consult, documentation, research) are not broken out on this site. And a benchmark score is not a clinical safety certification: it measures rubric adherence on 525 hard tasks, nothing more.

Updates

The board changes when a vendor publishes a new result, when a document is located for a row that was source pending, or when a higher-authority document replaces the one a row was read from. Every change, including price updates, is dated on the updates page, and each release of the results is published in full on the data page. The current table covers 31 models as of September 30, 2026.