Humanity's Last Exam leaderboard: which LLMs score highest
Humanity's Last Exam (expert-level questions across domains) currently has scores for 29 models. The top score is 40.6% by Muse Spark; the median model scores 10.7% and the lowest scores 2.7%.
| # | Model | Organisation | Humanity's Last Exam score | Cheapest input $/M | Source | Captured |
|---|---|---|---|---|---|---|
| 1 | Muse Spark | Meta AI | 40.6% | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 2 | Gemini 3 Pro Preview | Google DeepMind | 37.5% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 3 | GPT-5.4 (xhigh) | OpenAI | 36.2% | $2.50 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 4 | Claude Opus 4.6 (max) | Anthropic | 34.4% | $5.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 5 | GPT-5 Pro | OpenAI | 31.6% | $15.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 6 | GPT-5 (high) | OpenAI | 25.3% | $1.25 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 7 | Kimi K2.5 | Moonshot | 24.4% | $0.450 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 8 | Gemini 2.5 Pro Preview (Jun 2025) | Google DeepMind | 21.6% | $1.25 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 9 | o3 (high) | OpenAI | 20.3% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 10 | Gemini 2.5 Pro Exp (Mar 2025) | Google DeepMind | 18.2% | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 11 | o4-mini (high) | OpenAI | 18.1% | $1.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 12 | Gemini 2.5 Flash Preview (Apr 2025) | Google DeepMind | 12.1% | $0.300 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 13 | Claude Opus 4.1 (unknown thinking) | Anthropic | 11.5% | $15.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 14 | gemini-2.5-flash-preview-05-20 | Google DeepMind | 11.0% | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 15 | Claude Opus 4 | Anthropic | 10.7% | $15.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 16 | glm-4.5 | Z.ai (Zhipu AI),Tsinghua University | 8.3% | $0.400 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 17 | GLM-4.5-Air | Z.ai (Zhipu AI),Tsinghua University | 8.1% | $0.125 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 18 | o1 Pro | OpenAI | 8.1% | $150.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 19 | Claude Sonnet 4 (unknown thinking) | Anthropic | 7.8% | $3.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 20 | Gemini 2.0 Flash Thinking Exp | Google DeepMind,Google | 6.6% | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 21 | Llama 4 Maverick | Meta AI | 5.7% | $0.188 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 22 | GPT-4.5 Preview (Feb 2025) | OpenAI | 5.4% | $75.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 23 | GPT-4.1 | OpenAI | 5.4% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 24 | gemini-1.5-pro-002 | Google DeepMind | 4.6% | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 25 | mistral-medium-2505 | Mistral AI | 4.5% | $0.400 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 26 | amazon.nova-pro-v1:0 | Amazon | 4.4% | $0.800 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 27 | Claude 3.5 Sonnet (Oct 2024) | Anthropic | 4.1% | $3.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 28 | amazon.nova-lite-v1:0 | Amazon | 3.6% | $0.060 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 29 | GPT-4o (Nov 2024) | OpenAI | 2.7% | $2.50 | Epoch AI Benchmarking Hub | 2026-10-05 |
Compare all benchmarks side by side on the live leaderboard →
Among the ten highest scorers, the cheapest listed API price belongs to Kimi K2.5 at $0.450 per million input tokens (score 24.4%).
What Humanity's Last Exam measures
Humanity's Last Exam (HLE) is a collection of expert-written questions across many academic fields, built by the Center for AI Safety and Scale AI to stay hard as models improve.
How to read these scores
Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.
Where the data comes from
Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.