Best LLMs for Reasoning, Maths and Science (2026)

Claude Fable 5 (max) leads this ranking with a score of 59.7 (50 is the neutral midpoint). This page ranks models on reasoning, maths and science benchmarks: graduate-level science questions (GPQA Diamond), competition and research mathematics (AIME-style problems, FrontierMath), expert exam questions (Humanity's Last Exam) and abstract puzzles (ARC-AGI-2). Data last captured 2026-10-05.

Best LLMs for Reasoning and Maths: top 30
#ModelOrganisationScoreOverall rankCheapest input $/M
1Claude Fable 5 (max)Anthropic59.7#9$10.00
2Claude Opus 5.5 (max)Anthropic59.4#1$4.00
3GPT-6.1 Sol (max)OpenAI59.0#5$2.00
4GPT-6 Astra (max)OpenAI59.0#2$10.00
5Claude Fable 5.1 (max)Anthropic59.0#3$10.00
6Claude Sonnet 5.5 (max)Anthropic58.8#4$2.00
7GPT-5.6 Sol (max)OpenAI58.7#7$2.00
8Claude Opus 5 (max)Anthropic58.5#6$5.00
9GPT-6 Sol (max)OpenAI58.3#8$2.00
10GPT-5.4 (xhigh)OpenAI58.0#13$2.50
11GPT-5.5 (xhigh)OpenAI57.7#10$5.00
12GPT-5.6 Terra (max)OpenAI57.2#11$2.00
13GPT-5.5 Pro (xhigh)OpenAI56.9#12$30.00
14Gemini 3 Pro PreviewGoogle DeepMind56.4#30$2.00
15GPT-5.4 Pro (xhigh)OpenAI56.3#15$30.00
16GPT-5.6 Luna (max)OpenAI55.2#19$0.200
17Kimi K3 (max)Moonshot54.9#17$1.29
18GPT-6 Luna (max)OpenAI54.6#34$0.100
19Gemini 3.7 Flash (high)Google DeepMind54.1#29$0.750
20Gemini 3.8 Flash (high)Google DeepMind54.1#37$0.750
21Qwen 3.6 Max (Preview)Alibaba53.9#27—
22Qwen3.8 Max (xhigh)Alibaba53.8#39$1.65
23Grok 4.6 (xhigh)xAI53.7#36$1.25
24Claude Opus 4.6 (max)Anthropic53.5#67$5.00
25GPT-5 ProOpenAI53.5#103$15.00
26Muse SparkMeta AI53.4#32—
27Muse Spark 1.3 (max)Meta AI53.3#20$1.25
28grok-4-0709xAI53.3#40—
29Qwen3.7 MaxAlibaba53.1#69$1.25
30DeepSeek V4 Pro 0813 (max)DeepSeek53.1#25$0.660

How this ranking is built

This ranking uses only reasoning, maths and science benchmarks: graduate-level science questions (GPQA Diamond), competition and research mathematics (AIME-style problems, FrontierMath), expert exam questions (Humanity's Last Exam) and abstract puzzles (ARC-AGI-2). For each model we take its z-score on every one of these benchmarks it has been measured on and shrink the average toward 50 when it has few results; models need at least three of them to appear. Benchmarks used: GPQA Diamond, OTIS Mock AIME, Humanity's Last Exam, FrontierMath T1-3, FrontierMath T4, ARC-AGI-2, SimpleBench, CritPt, ProofBench, LiveBench Reasoning, LiveBench Mathematics, EnigmaEval and MMLU-Pro.

More rankings