Chess Puzzles leaderboard: which LLMs score highest

Chess Puzzles (chess tactics puzzles) currently has scores for 124 models. The top score is 72.0% by GPT-6 Astra (max); the median model scores 13.0% and the lowest scores 0.0%.

Chess Puzzles leaderboard: top 50 of 124 models (highest-effort row per model)
#ModelOrganisationChess Puzzles scoreCheapest input $/MSourceCaptured
1GPT-6 Astra (max)OpenAI72.0%$10.00Epoch AI Benchmarking Hub2026-10-05
2GPT-5.6 Sol (pro, max)OpenAI64.0%$2.00Epoch AI Benchmarking Hub2026-10-05
3Gemini 3.8 Flash (high)Google DeepMind61.0%$0.750Epoch AI Benchmarking Hub2026-10-05
4GPT-6.1 Sol (max)OpenAI61.0%$2.00Epoch AI Benchmarking Hub2026-10-05
5GPT-5.4 Pro (xhigh)OpenAI58.6%$30.00Epoch AI Benchmarking Hub2026-10-05
6GPT-5.6 Sol (max)OpenAI55.0%$2.00Epoch AI Benchmarking Hub2026-10-05
7GPT-5.6 Terra (max)OpenAI54.0%$2.00Epoch AI Benchmarking Hub2026-10-05
8GPT-5.2 (xhigh)OpenAI49.0%$1.75Epoch AI Benchmarking Hub2026-10-05
9Claude Fable 5.1 (max)Anthropic47.0%$10.00Epoch AI Benchmarking Hub2026-10-05
10DeepSeek V4 Pro 0813 (max)DeepSeek47.0%$0.660Epoch AI Benchmarking Hub2026-10-05
11Gemini 3.7 Flash (high)Google DeepMind47.0%$0.750Epoch AI Benchmarking Hub2026-10-05
12GPT-5.4 (xhigh)OpenAI44.0%$2.50Epoch AI Benchmarking Hub2026-10-05
13Claude Opus 5 (max)Anthropic42.0%$5.00Epoch AI Benchmarking Hub2026-10-05
14Claude Fable 5 (max)Anthropic41.0%$10.00Epoch AI Benchmarking Hub2026-10-05
15GPT-5.6 Luna (max)OpenAI40.0%$0.200Epoch AI Benchmarking Hub2026-10-05
16Qwen3.8 Max (0902) (xhigh)Alibaba40.0%$1.65Epoch AI Benchmarking Hub2026-10-05
17Kimi K3 (max)Moonshot39.0%$1.29Epoch AI Benchmarking Hub2026-10-05
18Grok 4.7 (xhigh)xAI38.0%$2.00Epoch AI Benchmarking Hub2026-10-05
19Muse Spark 1.3 (max)Meta AI38.0%$1.25Epoch AI Benchmarking Hub2026-10-05
20GPT-5 (high)OpenAI37.0%$1.25Epoch AI Benchmarking Hub2026-10-05
21Grok 4.5 (high)xAI36.0%$2.00Epoch AI Benchmarking Hub2026-10-05
22o3 (high)OpenAI34.0%$2.00Epoch AI Benchmarking Hub2026-10-05
23DeepSeek V4 Flash 0731 (max)DeepSeek33.0%$0.015Epoch AI Benchmarking Hub2026-10-05
24GPT-5.1 (high)OpenAI32.0%$1.25Epoch AI Benchmarking Hub2026-10-05
25Gemini 3 Pro PreviewGoogle DeepMind31.0%$2.00Epoch AI Benchmarking Hub2026-10-05
26GPT-6 Luna (max)OpenAI31.0%$0.100Epoch AI Benchmarking Hub2026-10-05
27Grok 4.6 (xhigh)xAI31.0%$1.25Epoch AI Benchmarking Hub2026-10-05
28GPT-5 mini (high)OpenAI30.0%$0.250Epoch AI Benchmarking Hub2026-10-05
29Qwen3.8 Max (xhigh)Alibaba29.0%$1.65Epoch AI Benchmarking Hub2026-10-05
30grok-4-0709xAI28.0%—Epoch AI Benchmarking Hub2026-10-05
31GPT-5 nano (high)OpenAI27.0%$0.050Epoch AI Benchmarking Hub2026-10-05
32Kimi K2.6Moonshot26.0%$0.650Epoch AI Benchmarking Hub2026-10-05
33o4-mini (high)OpenAI26.0%$1.00Epoch AI Benchmarking Hub2026-10-05
34Qwen3.6 35B-A3BAlibaba26.0%$0.050Epoch AI Benchmarking Hub2026-10-05
35grok-4.3 (high)xAI25.0%$1.25Epoch AI Benchmarking Hub2026-10-05
36GPT-5.4 mini (xhigh)OpenAI24.0%$0.750Epoch AI Benchmarking Hub2026-10-05
37grok-4.20-0309-reasoningxAI24.0%$1.25Epoch AI Benchmarking Hub2026-10-05
38Qwen3.7 PlusAlibaba24.0%$0.282Epoch AI Benchmarking Hub2026-10-05
39Qwen3.7 FlashAlibaba23.0%$0.030Epoch AI Benchmarking Hub2026-10-05
40Qwen3.5 PlusAlibaba22.0%—Epoch AI Benchmarking Hub2026-10-05
41Qwen3.6 27BAlibaba22.0%$0.150Epoch AI Benchmarking Hub2026-10-05
42GLM-5.3 (max)Z.ai (Zhipu AI)21.0%$0.070Epoch AI Benchmarking Hub2026-10-05
43Inkling (xhigh)Thinking Machines21.0%$0.950Epoch AI Benchmarking Hub2026-10-05
44Kimi K2.7 CodeMoonshot21.0%$0.671Epoch AI Benchmarking Hub2026-10-05
45Qwen3.5 FlashAlibaba21.0%$0.065Epoch AI Benchmarking Hub2026-10-05
46DeepSeek v4 Pro (max)DeepSeek20.0%$0.435Epoch AI Benchmarking Hub2026-10-05
47gpt-oss-120b (high)OpenAI20.0%$0.030Epoch AI Benchmarking Hub2026-10-05
48Kimi K2 Thinking TurboMoonshot20.0%—Epoch AI Benchmarking Hub2026-10-05
49Qwen 3.6 Flash (2026-04-16)Alibaba20.0%—Epoch AI Benchmarking Hub2026-10-05
50Qwen 3.6 Max (Preview)Alibaba20.0%—Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to DeepSeek V4 Pro 0813 (max) at $0.660 per million input tokens (score 47.0%).

What Chess Puzzles measures

Chess Puzzles measures chess tactics puzzles.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks