FrontierMath T4 leaderboard: which LLMs score highest

FrontierMath T4 (research-level maths, tier 4) currently has scores for 52 models. The top score is 100.0% by GPT-6.1 Sol (max); the median model scores 31.7% and the lowest scores 0.0%.

FrontierMath T4 leaderboard: top 50 of 52 models (highest-effort row per model)
#ModelOrganisationFrontierMath T4 scoreCheapest input $/MSourceCaptured
1GPT-6.1 Sol (max)OpenAI100.0%$2.00Epoch AI Benchmarking Hub2026-10-05
2GPT-6 Astra (max)OpenAI97.6%$10.00Epoch AI Benchmarking Hub2026-10-05
3Claude Opus 5.5 (max)Anthropic95.0%$4.00Epoch AI Benchmarking Hub2026-10-05
4Claude Fable 5 (max)Anthropic90.2%$10.00Epoch AI Benchmarking Hub2026-10-05
5GPT-6 Sol (max)OpenAI90.0%$2.00Epoch AI Benchmarking Hub2026-10-05
6Claude Fable 5.1 (max)Anthropic87.8%$10.00Epoch AI Benchmarking Hub2026-10-05
7GPT-5.6 Sol (max)OpenAI82.9%$2.00Epoch AI Benchmarking Hub2026-10-05
8Claude Sonnet 5.5 (max)Anthropic80.5%$2.00Epoch AI Benchmarking Hub2026-10-05
9GPT-5.6 Sol (pro, max)OpenAI80.5%$2.00Epoch AI Benchmarking Hub2026-10-05
10GPT-5.5 Pro (xhigh)OpenAI78.0%$30.00Epoch AI Benchmarking Hub2026-10-05
11AI co-mathematicianGoogle DeepMind75.6%—Epoch AI Benchmarking Hub2026-10-05
12Claude Opus 5 (max)Anthropic73.2%$5.00Epoch AI Benchmarking Hub2026-10-05
13GPT-5.5 (xhigh)OpenAI72.5%$5.00Epoch AI Benchmarking Hub2026-10-05
14GPT-5.6 Terra (max)OpenAI70.7%$2.00Epoch AI Benchmarking Hub2026-10-05
15GPT-5.6 Luna (max)OpenAI61.0%$0.200Epoch AI Benchmarking Hub2026-10-05
16GPT-5.4 Pro (xhigh)OpenAI58.5%$30.00Epoch AI Benchmarking Hub2026-10-05
17GPT-6 Luna (max)OpenAI56.1%$0.100Epoch AI Benchmarking Hub2026-10-05
18GPT-5.4 (xhigh)OpenAI49.0%$2.50Epoch AI Benchmarking Hub2026-10-05
19Muse Spark 1.3 (max)Meta AI46.3%$1.25Epoch AI Benchmarking Hub2026-10-05
20Qwen3.8 Max (xhigh)Alibaba46.3%$1.65Epoch AI Benchmarking Hub2026-10-05
21Kimi K3 (max)Moonshot39.0%$1.29Epoch AI Benchmarking Hub2026-10-05
22Gemini 3.7 Flash (high)Google DeepMind36.6%$0.750Epoch AI Benchmarking Hub2026-10-05
23Qwen3.7 MaxAlibaba34.1%$1.25Epoch AI Benchmarking Hub2026-10-05
24Qwen3.8 Max (0902) (xhigh)Alibaba34.1%$1.65Epoch AI Benchmarking Hub2026-10-05
25Claude Opus 4.7 (max)Anthropic31.7%$5.00Epoch AI Benchmarking Hub2026-10-05
26Grok 4.6 (xhigh)xAI31.7%$1.25Epoch AI Benchmarking Hub2026-10-05
27GPT-5.2 (xhigh)OpenAI31.7%$1.75Epoch AI Benchmarking Hub2026-10-05
28Claude Sonnet 5 (max)Anthropic29.3%$2.00Epoch AI Benchmarking Hub2026-10-05
29GLM-5.3 (max)Z.ai (Zhipu AI)29.3%$0.070Epoch AI Benchmarking Hub2026-10-05
30Claude Opus 4.6 (max)Anthropic26.8%$5.00Epoch AI Benchmarking Hub2026-10-05
31DeepSeek V4 Pro 0813 (max)DeepSeek26.8%$0.660Epoch AI Benchmarking Hub2026-10-05
32Kimi K2.6Moonshot25.6%$0.650Epoch AI Benchmarking Hub2026-10-05
33DeepSeek V4 Flash 0731 (max)DeepSeek24.4%$0.015Epoch AI Benchmarking Hub2026-10-05
34Grok 4.5 (high)xAI24.4%$2.00Epoch AI Benchmarking Hub2026-10-05
35Gemini 3.8 Flash (high)Google DeepMind22.0%$0.750Epoch AI Benchmarking Hub2026-10-05
36GPT-5 (high)OpenAI22.0%$1.25Epoch AI Benchmarking Hub2026-10-05
37GPT-5 ProOpenAI19.5%$15.00Epoch AI Benchmarking Hub2026-10-05
38GLM-5.3-Flash (max)Z.ai (Zhipu AI)17.1%$0.110Epoch AI Benchmarking Hub2026-10-05
39Grok 4.7 (xhigh)xAI17.1%$2.00Epoch AI Benchmarking Hub2026-10-05
40grok-4.20-0309-reasoningxAI17.1%$1.25Epoch AI Benchmarking Hub2026-10-05
41Inkling Small (xhigh)Thinking Machines17.1%$0.450Epoch AI Benchmarking Hub2026-10-05
42grok-4.3 (high)xAI14.6%$1.25Epoch AI Benchmarking Hub2026-10-05
43GPT-5 mini (high)OpenAI12.2%$0.250Epoch AI Benchmarking Hub2026-10-05
44Kimi K2.7 CodeMoonshot12.2%$0.671Epoch AI Benchmarking Hub2026-10-05
45GPT-5.4 mini (xhigh)OpenAI9.8%$0.750Epoch AI Benchmarking Hub2026-10-05
46Inkling (xhigh)Thinking Machines4.9%$0.950Epoch AI Benchmarking Hub2026-10-05
47o4-mini (high)OpenAI4.9%$1.00Epoch AI Benchmarking Hub2026-10-05
48claude-opus-4-1-20250805_32KAnthropic2.4%—Epoch AI Benchmarking Hub2026-10-05
49DeepSeek v4 Pro (max)DeepSeek2.4%$0.435Epoch AI Benchmarking Hub2026-10-05
50GPT-5 nano (high)OpenAI2.4%$0.050Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to GPT-6.1 Sol (max) at $2.00 per million input tokens (score 100.0%).

What FrontierMath T4 measures

FrontierMath Tier 4 is the smaller, hardest extension of Epoch AI's FrontierMath set of original research-level maths problems, whose answers are exact values that can be checked automatically.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks