ALE-Bench leaderboard: which LLMs score highest

ALE-Bench (AtCoder heuristic-optimisation contests (performance rating)) currently has scores for 78 models. The top score is 2,951 by GPT-6 Astra (max); the median model scores 755 and the lowest scores 138.

ALE-Bench leaderboard: top 50 of 78 models (highest-effort row per model)
#ModelOrganisationALE-Bench scoreCheapest input $/MSourceCaptured
1GPT-6 Astra (max)OpenAI2,951$10.00Epoch AI Benchmarking Hub2026-10-05
2GPT-6 Sol (max)OpenAI2,462$2.00Epoch AI Benchmarking Hub2026-10-05
3GPT-5.6 Sol (max)OpenAI2,177$2.00Epoch AI Benchmarking Hub2026-10-05
4GPT-5.6 Terra (max)OpenAI1,951$2.00Epoch AI Benchmarking Hub2026-10-05
5GPT-5.5 (xhigh)OpenAI1,943$5.00Epoch AI Benchmarking Hub2026-10-05
6GPT-5.6 Luna (max)OpenAI1,667$0.200Epoch AI Benchmarking Hub2026-10-05
7GPT-5.3 Codex (xhigh)OpenAI1,655$1.75Epoch AI Benchmarking Hub2026-10-05
8Kimi K3 (max)Moonshot1,524$1.29Epoch AI Benchmarking Hub2026-10-05
9Grok 4.6 (xhigh)xAI1,508$1.25Epoch AI Benchmarking Hub2026-10-05
10DeepSeek V4 Pro 0813 (max)DeepSeek1,403$0.660Epoch AI Benchmarking Hub2026-10-05
11Grok 4.5 (high)xAI1,309$2.00Epoch AI Benchmarking Hub2026-10-05
12DeepSeek V4 Flash 0731 (max)DeepSeek1,306$0.015Epoch AI Benchmarking Hub2026-10-05
13GPT-5.2 CodexOpenAI1,300$1.75Epoch AI Benchmarking Hub2026-10-05
14Gemini 3.8 Flash (high)Google DeepMind1,270$0.750Epoch AI Benchmarking Hub2026-10-05
15GPT-5.1 CodexOpenAI1,245$1.25Epoch AI Benchmarking Hub2026-10-05
16GPT-5.1-Codex-MaxOpenAI1,209$1.25Epoch AI Benchmarking Hub2026-10-05
17GPT-5.1 (high)OpenAI1,192$1.25Epoch AI Benchmarking Hub2026-10-05
18Qwen3.7 MaxAlibaba1,189$1.25Epoch AI Benchmarking Hub2026-10-05
19Gemini 3 Pro PreviewGoogle DeepMind1,177$2.00Epoch AI Benchmarking Hub2026-10-05
20GPT-5 (high)OpenAI1,162$1.25Epoch AI Benchmarking Hub2026-10-05
21grok-4-20xAI1,150$1.25Epoch AI Benchmarking Hub2026-10-05
22Kimi K2.6Moonshot1,093$0.650Epoch AI Benchmarking Hub2026-10-05
23deepseek-v4.1-flash-maxDeepSeek1,092—Epoch AI Benchmarking Hub2026-10-05
24o3 (high)OpenAI934$2.00Epoch AI Benchmarking Hub2026-10-05
25Gemma 4 26B A4BGoogle DeepMind927$0.076Epoch AI Benchmarking Hub2026-10-05
26Gemma 4 31B ITGoogle DeepMind926$0.100Epoch AI Benchmarking Hub2026-10-05
27Gemini 3.7 Flash (high)Google DeepMind904$0.750Epoch AI Benchmarking Hub2026-10-05
28mimo-v2.5-proXiaomi Corp900$0.435Epoch AI Benchmarking Hub2026-10-05
29GLM-5.1Z.ai (Zhipu AI)887$1.05Epoch AI Benchmarking Hub2026-10-05
30Kimi K2.7 CodeMoonshot886$0.671Epoch AI Benchmarking Hub2026-10-05
31o4-mini (high)OpenAI826$1.00Epoch AI Benchmarking Hub2026-10-05
32Kimi K2.5Moonshot822$0.450Epoch AI Benchmarking Hub2026-10-05
33DeepSeek-R1 (May 2025)DeepSeek804$0.400Epoch AI Benchmarking Hub2026-10-05
34GPT-5 mini (high)OpenAI800$0.250Epoch AI Benchmarking Hub2026-10-05
35Mercury 2Inception Labs786$0.250Epoch AI Benchmarking Hub2026-10-05
36gemini-2.5-pro_32KGoogle DeepMind786$1.25Epoch AI Benchmarking Hub2026-10-05
37mimo-v2-proXiaomi Corp785$1.10Epoch AI Benchmarking Hub2026-10-05
38GLM-5Z.ai (Zhipu AI)766$0.600Epoch AI Benchmarking Hub2026-10-05
39Mistral Medium 3.5Mistral AI764$1.50Epoch AI Benchmarking Hub2026-10-05
40DeepSeek-V3.1-TerminusDeepSeek745$0.270Epoch AI Benchmarking Hub2026-10-05
41MiMo-V2-FlashXiaomiMiMo738$0.110Epoch AI Benchmarking Hub2026-10-05
42GPT-5 nano (high)OpenAI719$0.050Epoch AI Benchmarking Hub2026-10-05
43Step 3.7 FlashStepFun694$0.200Epoch AI Benchmarking Hub2026-10-05
44claude-opus-4-1-20250805_16KAnthropic675—Epoch AI Benchmarking Hub2026-10-05
45Qwen 3.6 Plus (2026-04-02)Alibaba670—Epoch AI Benchmarking Hub2026-10-05
46gemini-2.5-flashGoogle DeepMind662$0.300Epoch AI Benchmarking Hub2026-10-05
47claude-sonnet-4-20250514_32KAnthropic655—Epoch AI Benchmarking Hub2026-10-05
48claude-haiku-4-5-20251001_32KAnthropic653$1.00Epoch AI Benchmarking Hub2026-10-05
49MiniMax-M3MiniMax640$0.230Epoch AI Benchmarking Hub2026-10-05
50MiniMax-M2.1MiniMax624$0.300Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to GPT-5.6 Luna (max) at $0.200 per million input tokens (score 1,667).

What ALE-Bench measures

ALE-Bench measures AtCoder heuristic-optimisation contests (performance rating).

How to read these scores

Scores are Elo-style ratings (higher is better). Only differences between models mean anything, and small gaps may sit inside the source's own margin of error. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks