Humanity's Last Exam leaderboard: which LLMs score highest

Humanity's Last Exam (expert-level questions across domains) currently has scores for 29 models. The top score is 40.6% by Muse Spark; the median model scores 10.7% and the lowest scores 2.7%.

Humanity's Last Exam leaderboard: top 29 of 29 models (highest-effort row per model)
#ModelOrganisationHumanity's Last Exam scoreCheapest input $/MSourceCaptured
1Muse SparkMeta AI40.6%—Epoch AI Benchmarking Hub2026-10-05
2Gemini 3 Pro PreviewGoogle DeepMind37.5%$2.00Epoch AI Benchmarking Hub2026-10-05
3GPT-5.4 (xhigh)OpenAI36.2%$2.50Epoch AI Benchmarking Hub2026-10-05
4Claude Opus 4.6 (max)Anthropic34.4%$5.00Epoch AI Benchmarking Hub2026-10-05
5GPT-5 ProOpenAI31.6%$15.00Epoch AI Benchmarking Hub2026-10-05
6GPT-5 (high)OpenAI25.3%$1.25Epoch AI Benchmarking Hub2026-10-05
7Kimi K2.5Moonshot24.4%$0.450Epoch AI Benchmarking Hub2026-10-05
8Gemini 2.5 Pro Preview (Jun 2025)Google DeepMind21.6%$1.25Epoch AI Benchmarking Hub2026-10-05
9o3 (high)OpenAI20.3%$2.00Epoch AI Benchmarking Hub2026-10-05
10Gemini 2.5 Pro Exp (Mar 2025)Google DeepMind18.2%—Epoch AI Benchmarking Hub2026-10-05
11o4-mini (high)OpenAI18.1%$1.00Epoch AI Benchmarking Hub2026-10-05
12Gemini 2.5 Flash Preview (Apr 2025)Google DeepMind12.1%$0.300Epoch AI Benchmarking Hub2026-10-05
13Claude Opus 4.1 (unknown thinking)Anthropic11.5%$15.00Epoch AI Benchmarking Hub2026-10-05
14gemini-2.5-flash-preview-05-20Google DeepMind11.0%—Epoch AI Benchmarking Hub2026-10-05
15Claude Opus 4Anthropic10.7%$15.00Epoch AI Benchmarking Hub2026-10-05
16glm-4.5Z.ai (Zhipu AI),Tsinghua University8.3%$0.400Epoch AI Benchmarking Hub2026-10-05
17GLM-4.5-AirZ.ai (Zhipu AI),Tsinghua University8.1%$0.125Epoch AI Benchmarking Hub2026-10-05
18o1 ProOpenAI8.1%$150.00Epoch AI Benchmarking Hub2026-10-05
19Claude Sonnet 4 (unknown thinking)Anthropic7.8%$3.00Epoch AI Benchmarking Hub2026-10-05
20Gemini 2.0 Flash Thinking ExpGoogle DeepMind,Google6.6%—Epoch AI Benchmarking Hub2026-10-05
21Llama 4 MaverickMeta AI5.7%$0.188Epoch AI Benchmarking Hub2026-10-05
22GPT-4.5 Preview (Feb 2025)OpenAI5.4%$75.00Epoch AI Benchmarking Hub2026-10-05
23GPT-4.1OpenAI5.4%$2.00Epoch AI Benchmarking Hub2026-10-05
24gemini-1.5-pro-002Google DeepMind4.6%—Epoch AI Benchmarking Hub2026-10-05
25mistral-medium-2505Mistral AI4.5%$0.400Epoch AI Benchmarking Hub2026-10-05
26amazon.nova-pro-v1:0Amazon4.4%$0.800Epoch AI Benchmarking Hub2026-10-05
27Claude 3.5 Sonnet (Oct 2024)Anthropic4.1%$3.00Epoch AI Benchmarking Hub2026-10-05
28amazon.nova-lite-v1:0Amazon3.6%$0.060Epoch AI Benchmarking Hub2026-10-05
29GPT-4o (Nov 2024)OpenAI2.7%$2.50Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to Kimi K2.5 at $0.450 per million input tokens (score 24.4%).

What Humanity's Last Exam measures

Humanity's Last Exam (HLE) is a collection of expert-written questions across many academic fields, built by the Center for AI Safety and Scale AI to stay hard as models improve.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks