EnigmaEval leaderboard: which LLMs score highest

EnigmaEval (multi-step puzzle hunts) currently has scores for 26 models. The top score is 18.8% by GPT-5 Pro; the median model scores 3.2% and the lowest scores 0.4%.

EnigmaEval leaderboard: top 26 of 26 models (highest-effort row per model)
#ModelOrganisationEnigmaEval scoreCheapest input $/MSourceCaptured
1GPT-5 ProOpenAI18.8%$15.00Epoch AI Benchmarking Hub2026-10-05
2Gemini 3 Pro PreviewGoogle DeepMind18.2%$2.00Epoch AI Benchmarking Hub2026-10-05
3GPT-5.4 (xhigh)OpenAI16.0%$2.50Epoch AI Benchmarking Hub2026-10-05
4o3 (high)OpenAI11.9%$2.00Epoch AI Benchmarking Hub2026-10-05
5o4-mini (high)OpenAI9.2%$1.00Epoch AI Benchmarking Hub2026-10-05
6Claude Opus 4.6 (max)Anthropic7.6%$5.00Epoch AI Benchmarking Hub2026-10-05
7Claude Opus 4.1 (unknown thinking)Anthropic7.2%$15.00Epoch AI Benchmarking Hub2026-10-05
8o1 ProOpenAI6.1%$150.00Epoch AI Benchmarking Hub2026-10-05
9Claude Opus 4Anthropic5.6%$15.00Epoch AI Benchmarking Hub2026-10-05
10Gemini 2.5 Pro Preview (Jun 2025)Google DeepMind5.6%$1.25Epoch AI Benchmarking Hub2026-10-05
11Gemini 2.5 Pro Exp (Mar 2025)Google DeepMind4.1%—Epoch AI Benchmarking Hub2026-10-05
12Kimi K2.5Moonshot3.4%$0.450Epoch AI Benchmarking Hub2026-10-05
13GPT-4.5 Preview (Feb 2025)OpenAI3.2%$75.00Epoch AI Benchmarking Hub2026-10-05
14Claude Sonnet 4 (unknown thinking)Anthropic3.1%$3.00Epoch AI Benchmarking Hub2026-10-05
15gemini-2.5-flash-preview-05-20Google DeepMind2.7%—Epoch AI Benchmarking Hub2026-10-05
16claude-3-7-sonnet-20250219Anthropic2.3%—Epoch AI Benchmarking Hub2026-10-05
17GPT-4.1OpenAI2.2%$2.00Epoch AI Benchmarking Hub2026-10-05
18Gemini 2.0 Flash Thinking ExpGoogle DeepMind,Google1.1%—Epoch AI Benchmarking Hub2026-10-05
19Claude 3.5 Sonnet (Oct 2024)Anthropic0.9%$3.00Epoch AI Benchmarking Hub2026-10-05
20pixtral-large-2411Mistral AI0.8%—Epoch AI Benchmarking Hub2026-10-05
21claude-3-opus-20240229Anthropic0.8%$15.00Epoch AI Benchmarking Hub2026-10-05
22GPT-4o (Nov 2024)OpenAI0.8%$2.50Epoch AI Benchmarking Hub2026-10-05
23Gemini 2.0 Pro Exp (Feb 2025)Google DeepMind0.7%—Epoch AI Benchmarking Hub2026-10-05
24gemini-2.0-flash-02-05Google DeepMind,Google0.6%—Epoch AI Benchmarking Hub2026-10-05
25Llama 4 MaverickMeta AI0.6%$0.188Epoch AI Benchmarking Hub2026-10-05
26Llama-3.2-90B-Vision-InstructMeta AI0.4%$2.00Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to o4-mini (high) at $1.00 per million input tokens (score 9.2%).

What EnigmaEval measures

EnigmaEval measures multi-step puzzle hunts.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks