OTIS Mock AIME leaderboard: which LLMs score highest

OTIS Mock AIME (competition maths (AIME 2024-25 style)) currently has scores for 188 models. The top score is 100.0% by Claude Fable 5.1 (max); the median model scores 57.8% and the lowest scores 0.0%.

OTIS Mock AIME leaderboard: top 50 of 188 models (highest-effort row per model)
#ModelOrganisationOTIS Mock AIME scoreCheapest input $/MSourceCaptured
1Claude Fable 5.1 (max)Anthropic100.0%$10.00Epoch AI Benchmarking Hub2026-10-05
2Claude Opus 5.5 (max)Anthropic100.0%$4.00Epoch AI Benchmarking Hub2026-10-05
3Claude Sonnet 5.5 (max)Anthropic100.0%$2.00Epoch AI Benchmarking Hub2026-10-05
4GPT-5.6 Sol (max)OpenAI100.0%$2.00Epoch AI Benchmarking Hub2026-10-05
5GPT-6 Astra (max)OpenAI100.0%$10.00Epoch AI Benchmarking Hub2026-10-05
6GPT-6 Sol (max)OpenAI100.0%$2.00Epoch AI Benchmarking Hub2026-10-05
7GPT-6.1 Sol (max)OpenAI100.0%$2.00Epoch AI Benchmarking Hub2026-10-05
8Qwen3.8 Max (0902) (xhigh)Alibaba100.0%$1.65Epoch AI Benchmarking Hub2026-10-05
9Claude Fable 5 (max)Anthropic99.7%$10.00Epoch AI Benchmarking Hub2026-10-05
10GPT-5.6 Terra (max)OpenAI99.7%$2.00Epoch AI Benchmarking Hub2026-10-05
11Qwen3.8 Max (xhigh)Alibaba99.4%$1.65Epoch AI Benchmarking Hub2026-10-05
12Grok 4.6 (xhigh)xAI99.2%$1.25Epoch AI Benchmarking Hub2026-10-05
13Claude Opus 5 (max)Anthropic98.9%$5.00Epoch AI Benchmarking Hub2026-10-05
14Gemini 3.8 Flash (high)Google DeepMind98.9%$0.750Epoch AI Benchmarking Hub2026-10-05
15GPT-6 Luna (max)OpenAI98.9%$0.100Epoch AI Benchmarking Hub2026-10-05
16DeepSeek V4 Pro 0813 (max)DeepSeek98.6%$0.660Epoch AI Benchmarking Hub2026-10-05
17GPT-5.6 Luna (max)OpenAI98.3%$0.200Epoch AI Benchmarking Hub2026-10-05
18Grok 4.7 (xhigh)xAI98.1%$2.00Epoch AI Benchmarking Hub2026-10-05
19Grok 4.5 (high)xAI97.8%$2.00Epoch AI Benchmarking Hub2026-10-05
20Gemini 3.7 Flash (high)Google DeepMind97.2%$0.750Epoch AI Benchmarking Hub2026-10-05
21Kimi K3 (max)Moonshot97.2%$1.29Epoch AI Benchmarking Hub2026-10-05
22DeepSeek v4 Pro (max)DeepSeek96.7%$0.435Epoch AI Benchmarking Hub2026-10-05
23Kimi K2.6Moonshot96.1%$0.650Epoch AI Benchmarking Hub2026-10-05
24GPT-5.2 (xhigh)OpenAI96.1%$1.75Epoch AI Benchmarking Hub2026-10-05
25Kimi K2.7 CodeMoonshot95.6%$0.671Epoch AI Benchmarking Hub2026-10-05
26Qwen3.7 MaxAlibaba95.6%$1.25Epoch AI Benchmarking Hub2026-10-05
27GPT-5.4 (xhigh)OpenAI95.3%$2.50Epoch AI Benchmarking Hub2026-10-05
28DeepSeek V4 Flash 0731 (max)DeepSeek94.4%$0.015Epoch AI Benchmarking Hub2026-10-05
29GLM-5.3-Flash (max)Z.ai (Zhipu AI)93.9%$0.110Epoch AI Benchmarking Hub2026-10-05
30GLM-5.1Z.ai (Zhipu AI)93.3%$1.05Epoch AI Benchmarking Hub2026-10-05
31grok-4.3 (high)xAI93.3%$1.25Epoch AI Benchmarking Hub2026-10-05
32Qwen 3.6 Plus (2026-04-02)Alibaba93.3%—Epoch AI Benchmarking Hub2026-10-05
33Qwen3.7 PlusAlibaba93.3%$0.282Epoch AI Benchmarking Hub2026-10-05
34grok-4.20-0309-reasoningxAI92.2%$1.25Epoch AI Benchmarking Hub2026-10-05
35Kimi K2.5 (Fireworks)Moonshot92.2%$0.450Epoch AI Benchmarking Hub2026-10-05
36Gemini 3 Pro PreviewGoogle DeepMind91.4%$2.00Epoch AI Benchmarking Hub2026-10-05
37GPT-5 (high)OpenAI91.4%$1.25Epoch AI Benchmarking Hub2026-10-05
38Claude Opus 4.6 (max)Anthropic91.1%$5.00Epoch AI Benchmarking Hub2026-10-05
39GLM-5.3 (max)Z.ai (Zhipu AI)91.1%$0.070Epoch AI Benchmarking Hub2026-10-05
40Qwen 3.6 Max (Preview)Alibaba91.1%—Epoch AI Benchmarking Hub2026-10-05
41Qwen3.6 27BAlibaba91.1%$0.150Epoch AI Benchmarking Hub2026-10-05
42Inkling Small (xhigh)Thinking Machines90.0%$0.450Epoch AI Benchmarking Hub2026-10-05
43Muse SparkMeta AI88.9%—Epoch AI Benchmarking Hub2026-10-05
44GPT-5.4 mini (xhigh)OpenAI88.9%$0.750Epoch AI Benchmarking Hub2026-10-05
45gpt-oss-120b (high)OpenAI88.9%$0.030Epoch AI Benchmarking Hub2026-10-05
46Inkling (xhigh)Thinking Machines88.9%$0.950Epoch AI Benchmarking Hub2026-10-05
47Qwen3.5 397B-A17BAlibaba88.9%$0.164Epoch AI Benchmarking Hub2026-10-05
48GPT-5.1 (high)OpenAI88.6%$1.25Epoch AI Benchmarking Hub2026-10-05
49Claude Opus 4.7 (max)Anthropic86.7%$5.00Epoch AI Benchmarking Hub2026-10-05
50GPT-5 mini (high)OpenAI86.7%$0.250Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to Qwen3.8 Max (0902) (xhigh) at $1.65 per million input tokens (score 100.0%).

What OTIS Mock AIME measures

This board uses mock exams written in the style of the American Invitational Mathematics Examination (AIME), a competition for strong high-school mathematicians in which every answer is an integer from 0 to 999.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks