Updates
Every change to the leaderboard, dated: models added, sources located or replaced, methodology corrections, and price corrections. Each entry keeps a stable anchor.
- 2026-09-07
Sourcing correction. The site previously described its scores as produced by Arcophos under a run protocol. That was wrong: Arcophos does not run HealthBench Professional. The scores are compiled from published documents and now come from the Arcophos benchmark results database. Every row carries a source field, links to its document on the sources page, and is marked source pending until a published document is matched to it. 0 of 31 rows are source pending as of this entry. The methodology page was rewritten to match. The JSON download gained a source object per row and the CSV gained source_name and source_url columns.
- 2026-08-16
Initial release. 31 models listed on the 525-task set: GPT-6 Astra (Anthropic run), Claude Sonnet 5.5, Claude Fable 5, Claude Opus 5.5, GPT-6 Astra, Claude Fable 5 (September card), Claude Fable 5.1, GPT-6 Luna, GPT-6 Sol, GPT-5.6 Sol, Claude Opus 5, Muse Spark 1.1, Claude Sonnet 5, GPT-5.6 Terra, Claude Opus 4.8 (Opus 4.8 grader), Grok 4.7, Claude Opus 4.8, GPT-5.6 Luna, Muse Spark, GPT-5.6 Sol (August), Claude Opus 4.7, GPT-5.5, Grok 4.6, GPT-5.4, GPT-5, GPT-5.2, Claude Sonnet 4.6, GPT-5.6 Luna (August), GPT-5.1, GPT-5.5 Instant, MAI-Thinking-1. List prices verified against vendor price sheets on this date.
The current table is on the leaderboard; versioned downloads are on the data page.