FrontierMath T1-3 leaderboard: which LLMs score highest

FrontierMath T1-3 (research-level maths, tiers 1-3) currently has scores for 69 models. The top score is 93.7% by GPT-6 Astra (max); the median model scores 55.8% and the lowest scores 0.0%.

FrontierMath T1-3 leaderboard: top 50 of 69 models (highest-effort row per model)
#ModelOrganisationFrontierMath T1-3 scoreCheapest input $/MSourceCaptured
1GPT-6 Astra (max)OpenAI93.7%$10.00Epoch AI Benchmarking Hub2026-10-05
2GPT-6.1 Sol (max)OpenAI93.7%$2.00Epoch AI Benchmarking Hub2026-10-05
3Claude Opus 5.5 (max)Anthropic91.2%$4.00Epoch AI Benchmarking Hub2026-10-05
4Claude Fable 5.1 (max)Anthropic90.2%$10.00Epoch AI Benchmarking Hub2026-10-05
5GPT-6 Sol (max)OpenAI89.8%$2.00Epoch AI Benchmarking Hub2026-10-05
6GPT-5.6 Sol (max)OpenAI89.1%$2.00Epoch AI Benchmarking Hub2026-10-05
7Claude Sonnet 5.5 (max)Anthropic88.8%$2.00Epoch AI Benchmarking Hub2026-10-05
8GPT-5.5 Pro (xhigh)OpenAI87.7%$30.00Epoch AI Benchmarking Hub2026-10-05
9Claude Fable 5 (max)Anthropic87.0%$10.00Epoch AI Benchmarking Hub2026-10-05
10GPT-5.6 Terra (max)OpenAI86.0%$2.00Epoch AI Benchmarking Hub2026-10-05
11Claude Opus 5 (max)Anthropic85.6%$5.00Epoch AI Benchmarking Hub2026-10-05
12GPT-5.5 (xhigh)OpenAI85.3%$5.00Epoch AI Benchmarking Hub2026-10-05
13GPT-5.4 Pro (xhigh)OpenAI82.5%$30.00Epoch AI Benchmarking Hub2026-10-05
14GPT-5.6 Luna (max)OpenAI82.1%$0.200Epoch AI Benchmarking Hub2026-10-05
15GPT-6 Luna (max)OpenAI78.9%$0.100Epoch AI Benchmarking Hub2026-10-05
16GPT-5.4 (xhigh)OpenAI78.6%$2.50Epoch AI Benchmarking Hub2026-10-05
17Qwen3.8 Max (xhigh)Alibaba74.7%$1.65Epoch AI Benchmarking Hub2026-10-05
18Muse Spark 1.3 (max)Meta AI74.0%$1.25Epoch AI Benchmarking Hub2026-10-05
19Kimi K3 (max)Moonshot72.2%$1.29Epoch AI Benchmarking Hub2026-10-05
20Gemini 3.7 Flash (high)Google DeepMind71.6%$0.750Epoch AI Benchmarking Hub2026-10-05
21Claude Opus 4.7 (max)Anthropic70.2%$5.00Epoch AI Benchmarking Hub2026-10-05
22GLM-5.3 (max)Z.ai (Zhipu AI)68.8%$0.070Epoch AI Benchmarking Hub2026-10-05
23Gemini 3.8 Flash (high)Google DeepMind68.4%$0.750Epoch AI Benchmarking Hub2026-10-05
24GPT-5.2 (xhigh)OpenAI67.4%$1.75Epoch AI Benchmarking Hub2026-10-05
25Claude Opus 4.6 (max)Anthropic66.0%$5.00Epoch AI Benchmarking Hub2026-10-05
26Grok 4.6 (xhigh)xAI66.0%$1.25Epoch AI Benchmarking Hub2026-10-05
27Claude Sonnet 5 (max)Anthropic65.6%$2.00Epoch AI Benchmarking Hub2026-10-05
28Qwen3.8 Max (0902) (xhigh)Alibaba65.6%$1.65Epoch AI Benchmarking Hub2026-10-05
29DeepSeek V4 Pro 0813 (max)DeepSeek64.6%$0.660Epoch AI Benchmarking Hub2026-10-05
30Qwen3.7 MaxAlibaba64.6%$1.25Epoch AI Benchmarking Hub2026-10-05
31DeepSeek V4 Flash 0731 (max)DeepSeek57.5%$0.015Epoch AI Benchmarking Hub2026-10-05
32Grok 4.5 (high)xAI57.2%$2.00Epoch AI Benchmarking Hub2026-10-05
33Kimi K2.6Moonshot57.2%$0.650Epoch AI Benchmarking Hub2026-10-05
34GLM-5.3-Flash (max)Z.ai (Zhipu AI)55.8%$0.110Epoch AI Benchmarking Hub2026-10-05
35GPT-5 ProOpenAI55.8%$15.00Epoch AI Benchmarking Hub2026-10-05
36GPT-5 (high)OpenAI55.4%$1.25Epoch AI Benchmarking Hub2026-10-05
37Kimi K2.7 CodeMoonshot54.0%$0.671Epoch AI Benchmarking Hub2026-10-05
38Grok 4.7 (xhigh)xAI53.0%$2.00Epoch AI Benchmarking Hub2026-10-05
39GPT-5.4 mini (xhigh)OpenAI51.2%$0.750Epoch AI Benchmarking Hub2026-10-05
40GPT-5 mini (high)OpenAI46.7%$0.250Epoch AI Benchmarking Hub2026-10-05
41Inkling Small (xhigh)Thinking Machines46.3%$0.450Epoch AI Benchmarking Hub2026-10-05
42DeepSeek v4 Pro (max)DeepSeek45.3%$0.435Epoch AI Benchmarking Hub2026-10-05
43grok-4.20-0309-reasoningxAI44.9%$1.25Epoch AI Benchmarking Hub2026-10-05
44grok-4.3 (high)xAI42.8%$1.25Epoch AI Benchmarking Hub2026-10-05
45Qwen 3.6 Plus (2026-04-02)Alibaba38.2%—Epoch AI Benchmarking Hub2026-10-05
46GLM-5.1Z.ai (Zhipu AI)36.8%$1.05Epoch AI Benchmarking Hub2026-10-05
47o4-mini (high)OpenAI36.1%$1.00Epoch AI Benchmarking Hub2026-10-05
48Qwen3.6 27BAlibaba35.1%$0.150Epoch AI Benchmarking Hub2026-10-05
49Qwen3.7 PlusAlibaba34.4%$0.282Epoch AI Benchmarking Hub2026-10-05
50Inkling (xhigh)Thinking Machines33.3%$0.950Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to GPT-6.1 Sol (max) at $2.00 per million input tokens (score 93.7%).

What FrontierMath T1-3 measures

FrontierMath is a set of original, unpublished research-level maths problems written with professional mathematicians by Epoch AI. Answers are exact values that can be checked automatically. Tiers 1-3 are the main set.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks