ProofBench leaderboard: which LLMs score highest

ProofBench (writing mathematical proofs) currently has scores for 43 models. The top score is 100.0% by Claude Fable 5.1 (max); the median model scores 20.0% and the lowest scores 0.0%.

ProofBench leaderboard: top 43 of 43 models (highest-effort row per model)
#ModelOrganisationProofBench scoreCheapest input $/MSourceCaptured
1Claude Fable 5.1 (max)Anthropic100.0%$10.00Epoch AI Benchmarking Hub2026-10-05
2Claude Opus 5.5 (max)Anthropic100.0%$4.00Epoch AI Benchmarking Hub2026-10-05
3Claude Sonnet 5.5 (max)Anthropic100.0%$2.00Epoch AI Benchmarking Hub2026-10-05
4Claude Opus 5 (max)Anthropic99.0%$5.00Epoch AI Benchmarking Hub2026-10-05
5Claude Fable 5 (max)Anthropic95.0%$10.00Epoch AI Benchmarking Hub2026-10-05
6GPT-5.6 Sol (max)OpenAI83.0%$2.00Epoch AI Benchmarking Hub2026-10-05
7Claude Sonnet 5 (max)Anthropic77.0%$2.00Epoch AI Benchmarking Hub2026-10-05
8Tencent Hy4 preview (unknown)Tencent75.0%—Epoch AI Benchmarking Hub2026-10-05
9GPT-5.6 Luna (max)OpenAI60.0%$0.200Epoch AI Benchmarking Hub2026-10-05
10GPT-5.4 (xhigh)OpenAI56.0%$2.50Epoch AI Benchmarking Hub2026-10-05
11Claude Opus 4.7 (max)Anthropic54.0%$5.00Epoch AI Benchmarking Hub2026-10-05
12Claude Opus 4.6 (max)Anthropic50.0%$5.00Epoch AI Benchmarking Hub2026-10-05
13GPT-5.5 (xhigh)OpenAI50.0%$5.00Epoch AI Benchmarking Hub2026-10-05
14GLM-5.3 (max)Z.ai (Zhipu AI)49.0%$0.070Epoch AI Benchmarking Hub2026-10-05
15Claude Sonnet 4.6 (max)Anthropic45.0%$3.00Epoch AI Benchmarking Hub2026-10-05
16Grok 4.5 (high)xAI31.0%$2.00Epoch AI Benchmarking Hub2026-10-05
17Qwen3.7 MaxAlibaba26.0%$1.25Epoch AI Benchmarking Hub2026-10-05
18GLM-5.1Z.ai (Zhipu AI)22.2%$1.05Epoch AI Benchmarking Hub2026-10-05
19mimo-v2.5-proXiaomi Corp22.0%$0.435Epoch AI Benchmarking Hub2026-10-05
20GLM-5.3-Flash (max)Z.ai (Zhipu AI)21.0%$0.110Epoch AI Benchmarking Hub2026-10-05
21GPT-5.4 mini (xhigh)OpenAI21.0%$0.750Epoch AI Benchmarking Hub2026-10-05
22Gemini 3 Pro PreviewGoogle DeepMind20.0%$2.00Epoch AI Benchmarking Hub2026-10-05
23GPT-5 (high)OpenAI18.0%$1.25Epoch AI Benchmarking Hub2026-10-05
24MiniMax-M3MiniMax18.0%$0.230Epoch AI Benchmarking Hub2026-10-05
25Muse SparkMeta AI17.0%—Epoch AI Benchmarking Hub2026-10-05
26DeepSeek v4 Pro (max)DeepSeek16.0%$0.435Epoch AI Benchmarking Hub2026-10-05
27Kimi K2.6Moonshot16.0%$0.650Epoch AI Benchmarking Hub2026-10-05
28mimo-v2.5Xiaomi Corp16.0%$0.140Epoch AI Benchmarking Hub2026-10-05
29GPT-5.2 (xhigh)OpenAI15.0%$1.75Epoch AI Benchmarking Hub2026-10-05
30grok-4.20-0309-reasoningxAI14.0%$1.25Epoch AI Benchmarking Hub2026-10-05
31GPT-5 nano (high)OpenAI12.0%$0.050Epoch AI Benchmarking Hub2026-10-05
32grok-4.3 (high)xAI11.0%$1.25Epoch AI Benchmarking Hub2026-10-05
33GPT-5 mini (high)OpenAI9.0%$0.250Epoch AI Benchmarking Hub2026-10-05
34GPT-5.1-Codex-MaxOpenAI9.0%$1.25Epoch AI Benchmarking Hub2026-10-05
35Mistral Medium 3.5Mistral AI9.0%$1.50Epoch AI Benchmarking Hub2026-10-05
36DeepSeek-V3.2 (Thinking; Fireworks)DeepSeek8.0%$0.259Epoch AI Benchmarking Hub2026-10-05
37GLM-4.7Z.ai (Zhipu AI)6.0%$0.400Epoch AI Benchmarking Hub2026-10-05
38grok-4-1-fast-reasoningxAI4.0%$0.200Epoch AI Benchmarking Hub2026-10-05
39MiniMax-M2.5MiniMax4.0%$0.270Epoch AI Benchmarking Hub2026-10-05
40MiniMax-M2.7MiniMax3.0%$0.210Epoch AI Benchmarking Hub2026-10-05
41nemotron-3-ultraNvidia2.0%$0.500Epoch AI Benchmarking Hub2026-10-05
42Laguna M.1Poolside0.0%—Epoch AI Benchmarking Hub2026-10-05
43Laguna XS.2Poolside0.0%—Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to GPT-5.6 Luna (max) at $0.200 per million input tokens (score 60.0%).

What ProofBench measures

ProofBench measures writing mathematical proofs.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks