EnigmaEval leaderboard: which LLMs score highest
EnigmaEval (multi-step puzzle hunts) currently has scores for 26 models. The top score is 18.8% by GPT-5 Pro; the median model scores 3.2% and the lowest scores 0.4%.
| # | Model | Organisation | EnigmaEval score | Cheapest input $/M | Source | Captured |
|---|---|---|---|---|---|---|
| 1 | GPT-5 Pro | OpenAI | 18.8% | $15.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 2 | Gemini 3 Pro Preview | Google DeepMind | 18.2% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 3 | GPT-5.4 (xhigh) | OpenAI | 16.0% | $2.50 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 4 | o3 (high) | OpenAI | 11.9% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 5 | o4-mini (high) | OpenAI | 9.2% | $1.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 6 | Claude Opus 4.6 (max) | Anthropic | 7.6% | $5.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 7 | Claude Opus 4.1 (unknown thinking) | Anthropic | 7.2% | $15.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 8 | o1 Pro | OpenAI | 6.1% | $150.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 9 | Claude Opus 4 | Anthropic | 5.6% | $15.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 10 | Gemini 2.5 Pro Preview (Jun 2025) | Google DeepMind | 5.6% | $1.25 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 11 | Gemini 2.5 Pro Exp (Mar 2025) | Google DeepMind | 4.1% | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 12 | Kimi K2.5 | Moonshot | 3.4% | $0.450 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 13 | GPT-4.5 Preview (Feb 2025) | OpenAI | 3.2% | $75.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 14 | Claude Sonnet 4 (unknown thinking) | Anthropic | 3.1% | $3.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 15 | gemini-2.5-flash-preview-05-20 | Google DeepMind | 2.7% | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 16 | claude-3-7-sonnet-20250219 | Anthropic | 2.3% | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 17 | GPT-4.1 | OpenAI | 2.2% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 18 | Gemini 2.0 Flash Thinking Exp | Google DeepMind,Google | 1.1% | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 19 | Claude 3.5 Sonnet (Oct 2024) | Anthropic | 0.9% | $3.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 20 | pixtral-large-2411 | Mistral AI | 0.8% | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 21 | claude-3-opus-20240229 | Anthropic | 0.8% | $15.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 22 | GPT-4o (Nov 2024) | OpenAI | 0.8% | $2.50 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 23 | Gemini 2.0 Pro Exp (Feb 2025) | Google DeepMind | 0.7% | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 24 | gemini-2.0-flash-02-05 | Google DeepMind,Google | 0.6% | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 25 | Llama 4 Maverick | Meta AI | 0.6% | $0.188 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 26 | Llama-3.2-90B-Vision-Instruct | Meta AI | 0.4% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
Compare all benchmarks side by side on the live leaderboard →
Among the ten highest scorers, the cheapest listed API price belongs to o4-mini (high) at $1.00 per million input tokens (score 9.2%).
What EnigmaEval measures
EnigmaEval measures multi-step puzzle hunts.
How to read these scores
Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.
Where the data comes from
Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.