Best LLMs for Reasoning, Maths and Science (2026)
Claude Fable 5 (max) leads this ranking with a score of 59.7 (50 is the neutral midpoint). This page ranks models on reasoning, maths and science benchmarks: graduate-level science questions (GPQA Diamond), competition and research mathematics (AIME-style problems, FrontierMath), expert exam questions (Humanity's Last Exam) and abstract puzzles (ARC-AGI-2). Data last captured 2026-10-05.
| # | Model | Organisation | Score | Overall rank | Cheapest input $/M |
|---|---|---|---|---|---|
| 1 | Claude Fable 5 (max) | Anthropic | 59.7 | #9 | $10.00 |
| 2 | Claude Opus 5.5 (max) | Anthropic | 59.4 | #1 | $4.00 |
| 3 | GPT-6.1 Sol (max) | OpenAI | 59.0 | #5 | $2.00 |
| 4 | GPT-6 Astra (max) | OpenAI | 59.0 | #2 | $10.00 |
| 5 | Claude Fable 5.1 (max) | Anthropic | 59.0 | #3 | $10.00 |
| 6 | Claude Sonnet 5.5 (max) | Anthropic | 58.8 | #4 | $2.00 |
| 7 | GPT-5.6 Sol (max) | OpenAI | 58.7 | #7 | $2.00 |
| 8 | Claude Opus 5 (max) | Anthropic | 58.5 | #6 | $5.00 |
| 9 | GPT-6 Sol (max) | OpenAI | 58.3 | #8 | $2.00 |
| 10 | GPT-5.4 (xhigh) | OpenAI | 58.0 | #13 | $2.50 |
| 11 | GPT-5.5 (xhigh) | OpenAI | 57.7 | #10 | $5.00 |
| 12 | GPT-5.6 Terra (max) | OpenAI | 57.2 | #11 | $2.00 |
| 13 | GPT-5.5 Pro (xhigh) | OpenAI | 56.9 | #12 | $30.00 |
| 14 | Gemini 3 Pro Preview | Google DeepMind | 56.4 | #30 | $2.00 |
| 15 | GPT-5.4 Pro (xhigh) | OpenAI | 56.3 | #15 | $30.00 |
| 16 | GPT-5.6 Luna (max) | OpenAI | 55.2 | #19 | $0.200 |
| 17 | Kimi K3 (max) | Moonshot | 54.9 | #17 | $1.29 |
| 18 | GPT-6 Luna (max) | OpenAI | 54.6 | #34 | $0.100 |
| 19 | Gemini 3.7 Flash (high) | Google DeepMind | 54.1 | #29 | $0.750 |
| 20 | Gemini 3.8 Flash (high) | Google DeepMind | 54.1 | #37 | $0.750 |
| 21 | Qwen 3.6 Max (Preview) | Alibaba | 53.9 | #27 | — |
| 22 | Qwen3.8 Max (xhigh) | Alibaba | 53.8 | #39 | $1.65 |
| 23 | Grok 4.6 (xhigh) | xAI | 53.7 | #36 | $1.25 |
| 24 | Claude Opus 4.6 (max) | Anthropic | 53.5 | #67 | $5.00 |
| 25 | GPT-5 Pro | OpenAI | 53.5 | #103 | $15.00 |
| 26 | Muse Spark | Meta AI | 53.4 | #32 | — |
| 27 | Muse Spark 1.3 (max) | Meta AI | 53.3 | #20 | $1.25 |
| 28 | grok-4-0709 | xAI | 53.3 | #40 | — |
| 29 | Qwen3.7 Max | Alibaba | 53.1 | #69 | $1.25 |
| 30 | DeepSeek V4 Pro 0813 (max) | DeepSeek | 53.1 | #25 | $0.660 |
How this ranking is built
This ranking uses only reasoning, maths and science benchmarks: graduate-level science questions (GPQA Diamond), competition and research mathematics (AIME-style problems, FrontierMath), expert exam questions (Humanity's Last Exam) and abstract puzzles (ARC-AGI-2). For each model we take its z-score on every one of these benchmarks it has been measured on and shrink the average toward 50 when it has few results; models need at least three of them to appear. Benchmarks used: GPQA Diamond, OTIS Mock AIME, Humanity's Last Exam, FrontierMath T1-3, FrontierMath T4, ARC-AGI-2, SimpleBench, CritPt, ProofBench, LiveBench Reasoning, LiveBench Mathematics, EnigmaEval and MMLU-Pro.