Mystery Games leaderboard: which LLMs score highest

Mystery Games (deductive reasoning puzzles) currently has scores for 57 models. The top score is 84.0% by GPT-6 Astra (max); the median model scores 22.0% and the lowest scores 0.0%.

Mystery Games leaderboard: top 50 of 57 models (highest-effort row per model)
#ModelOrganisationMystery Games scoreCheapest input $/MSourceCaptured
1GPT-6 Astra (max)OpenAI84.0%$10.00Epoch AI Benchmarking Hub2026-10-05
2GPT-6.1 Sol (max)OpenAI80.0%$2.00Epoch AI Benchmarking Hub2026-10-05
3Claude Opus 5.5 (max)Anthropic71.0%$4.00Epoch AI Benchmarking Hub2026-10-05
4Claude Sonnet 5.5 (max)Anthropic65.0%$2.00Epoch AI Benchmarking Hub2026-10-05
5Claude Opus 5 (max)Anthropic59.0%$5.00Epoch AI Benchmarking Hub2026-10-05
6Claude Fable 5.1 (max)Anthropic58.0%$10.00Epoch AI Benchmarking Hub2026-10-05
7GPT-5.6 Sol (max)OpenAI58.0%$2.00Epoch AI Benchmarking Hub2026-10-05
8GPT-5.5 (xhigh)OpenAI56.0%$5.00Epoch AI Benchmarking Hub2026-10-05
9GPT-6 Sol (max)OpenAI56.0%$2.00Epoch AI Benchmarking Hub2026-10-05
10Claude Fable 5 (max)Anthropic52.0%$10.00Epoch AI Benchmarking Hub2026-10-05
11Gemini 3.8 Flash (high)Google DeepMind47.0%$0.750Epoch AI Benchmarking Hub2026-10-05
12DeepSeek V4 Pro 0813 (max)DeepSeek43.0%$0.660Epoch AI Benchmarking Hub2026-10-05
13Qwen3.8 Max (xhigh)Alibaba38.0%$1.65Epoch AI Benchmarking Hub2026-10-05
14Gemini 3.7 Flash (high)Google DeepMind37.0%$0.750Epoch AI Benchmarking Hub2026-10-05
15GPT-5.4 (xhigh)OpenAI37.0%$2.50Epoch AI Benchmarking Hub2026-10-05
16Claude Sonnet 5 (max)Anthropic35.0%$2.00Epoch AI Benchmarking Hub2026-10-05
17GPT-5.6 Terra (max)OpenAI35.0%$2.00Epoch AI Benchmarking Hub2026-10-05
18DeepSeek V4 Flash 0731 (max)DeepSeek34.0%$0.015Epoch AI Benchmarking Hub2026-10-05
19Grok 4.6 (xhigh)xAI34.0%$1.25Epoch AI Benchmarking Hub2026-10-05
20GLM-5.3 (max)Z.ai (Zhipu AI)33.0%$0.070Epoch AI Benchmarking Hub2026-10-05
21Qwen3.7 MaxAlibaba32.0%$1.25Epoch AI Benchmarking Hub2026-10-05
22Grok 4.7 (xhigh)xAI29.0%$2.00Epoch AI Benchmarking Hub2026-10-05
23o3 (high)OpenAI29.0%$2.00Epoch AI Benchmarking Hub2026-10-05
24Claude Opus 4.7 (max)Anthropic28.0%$5.00Epoch AI Benchmarking Hub2026-10-05
25Kimi K3 (max)Moonshot26.0%$1.29Epoch AI Benchmarking Hub2026-10-05
26Claude Opus 4.6 (max)Anthropic25.0%$5.00Epoch AI Benchmarking Hub2026-10-05
27Muse Spark 1.3 (max)Meta AI25.0%$1.25Epoch AI Benchmarking Hub2026-10-05
28GPT-5 (high)OpenAI23.0%$1.25Epoch AI Benchmarking Hub2026-10-05
29Qwen3.6 35B-A3BAlibaba22.0%$0.050Epoch AI Benchmarking Hub2026-10-05
30claude-opus-4-1-20250805_24KAnthropic21.0%—Epoch AI Benchmarking Hub2026-10-05
31GPT-5.6 Luna (max)OpenAI21.0%$0.200Epoch AI Benchmarking Hub2026-10-05
32nemotron-3-ultraNvidia20.0%$0.500Epoch AI Benchmarking Hub2026-10-05
33Qwen3.5 FlashAlibaba20.0%$0.065Epoch AI Benchmarking Hub2026-10-05
34Qwen 3.6 Max (Preview)Alibaba19.0%—Epoch AI Benchmarking Hub2026-10-05
35Kimi K2.6Moonshot18.0%$0.650Epoch AI Benchmarking Hub2026-10-05
36Qwen 3.6 Flash (2026-04-16)Alibaba18.0%—Epoch AI Benchmarking Hub2026-10-05
37Qwen3.5 397B-A17BAlibaba18.0%$0.164Epoch AI Benchmarking Hub2026-10-05
38Qwen3.5 PlusAlibaba17.0%—Epoch AI Benchmarking Hub2026-10-05
39Qwen3.7 PlusAlibaba17.0%$0.282Epoch AI Benchmarking Hub2026-10-05
40Qwen3.7 FlashAlibaba15.0%$0.030Epoch AI Benchmarking Hub2026-10-05
41GPT-4 (Jun 2023)OpenAI12.0%$30.00Epoch AI Benchmarking Hub2026-10-05
42gpt-4o-mini-2024-07-18OpenAI12.0%$0.150Epoch AI Benchmarking Hub2026-10-05
43Qwen 3.6 Plus (2026-04-02)Alibaba12.0%—Epoch AI Benchmarking Hub2026-10-05
44Qwen3-235B-A22B-Thinking-2507Alibaba9.0%$0.230Epoch AI Benchmarking Hub2026-10-05
45GLM-5.3-Flash (max)Z.ai (Zhipu AI)8.0%$0.110Epoch AI Benchmarking Hub2026-10-05
46GPT-5 nano (high)OpenAI8.0%$0.050Epoch AI Benchmarking Hub2026-10-05
47GPT-4.1 miniOpenAI7.0%$0.400Epoch AI Benchmarking Hub2026-10-05
48GPT-6 Luna (max)OpenAI7.0%$0.100Epoch AI Benchmarking Hub2026-10-05
49MiniMax-M3MiniMax7.0%$0.230Epoch AI Benchmarking Hub2026-10-05
50o3-mini (high)OpenAI7.0%$1.10Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to GPT-6.1 Sol (max) at $2.00 per million input tokens (score 80.0%).

What Mystery Games measures

Mystery Games measures deductive reasoning puzzles.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks