BALROG leaderboard: which LLMs score highest

BALROG (long-horizon game-playing agents) currently has scores for 34 models. The top score is 68.3% by GPT-6 Astra (max); the median model scores 27.6% and the lowest scores 3.7%.

BALROG leaderboard: top 34 of 34 models (highest-effort row per model)
#ModelOrganisationBALROG scoreCheapest input $/MSourceCaptured
1GPT-6 Astra (max)OpenAI68.3%$10.00Epoch AI Benchmarking Hub2026-10-05
2Claude Opus 5 (max)Anthropic63.4%$5.00Epoch AI Benchmarking Hub2026-10-05
3GPT-5.6 Sol (max)OpenAI60.0%$2.00Epoch AI Benchmarking Hub2026-10-05
4Gemini 3 Pro PreviewGoogle DeepMind58.1%$2.00Epoch AI Benchmarking Hub2026-10-05
5GPT-5.6 Terra (max)OpenAI53.2%$2.00Epoch AI Benchmarking Hub2026-10-05
6GPT-5.6 Luna (max)OpenAI45.6%$0.200Epoch AI Benchmarking Hub2026-10-05
7grok-4-0709xAI43.6%—Epoch AI Benchmarking Hub2026-10-05
8Gemini 2.5 Pro Exp (Mar 2025)Google DeepMind43.3%—Epoch AI Benchmarking Hub2026-10-05
9DeepSeek-R1DeepSeek34.9%$0.400Epoch AI Benchmarking Hub2026-10-05
10gemini-2.5-flashGoogle DeepMind33.5%$0.300Epoch AI Benchmarking Hub2026-10-05
11Claude 3.5 Sonnet (Oct 2024)Anthropic32.6%$3.00Epoch AI Benchmarking Hub2026-10-05
12GPT-4o (May 2024)OpenAI32.3%$2.50Epoch AI Benchmarking Hub2026-10-05
13claude-haiku-4-5-20251001Anthropic31.2%$1.00Epoch AI Benchmarking Hub2026-10-05
14claude-haiku-4-5-20251001_1KAnthropic31.2%—Epoch AI Benchmarking Hub2026-10-05
15grok-3-betaxAI29.5%—Epoch AI Benchmarking Hub2026-10-05
16Reka Flash 3rekaai29.2%$0.100Epoch AI Benchmarking Hub2026-10-05
17Llama-3.1-70B-InstructMeta AI27.9%$0.400Epoch AI Benchmarking Hub2026-10-05
18Llama-3.2-90B-Vision-InstructMeta AI27.3%$2.00Epoch AI Benchmarking Hub2026-10-05
19Llama-3.3-70B-InstructMeta AI23.0%$0.120Epoch AI Benchmarking Hub2026-10-05
20gemini-1.5-pro-002Google DeepMind21.0%—Epoch AI Benchmarking Hub2026-10-05
21DeepSeek-R1-Distill-Qwen-32BDeepSeek19.5%$0.150Epoch AI Benchmarking Hub2026-10-05
22Claude 3.5 Haiku (Oct 2024)Anthropic19.3%$0.800Epoch AI Benchmarking Hub2026-10-05
23Mistral-Nemo-Instruct-2407Mistral AI17.6%$0.019Epoch AI Benchmarking Hub2026-10-05
24gpt-4o-mini-2024-07-18OpenAI17.4%$0.150Epoch AI Benchmarking Hub2026-10-05
25Llama-3.2-11B-Vision-InstructMeta AI16.8%$0.049Epoch AI Benchmarking Hub2026-10-05
26qwen2.5-72b-instructAlibaba16.2%$0.120Epoch AI Benchmarking Hub2026-10-05
27Llama-3.1-8B-InstructMeta AI15.1%$0.020Epoch AI Benchmarking Hub2026-10-05
28gemini-1.5-flash-002Google DeepMind14.6%—Epoch AI Benchmarking Hub2026-10-05
29Qwen2-VL-72B-InstructAlibaba12.8%$0.130Epoch AI Benchmarking Hub2026-10-05
30phi-4Microsoft Research11.6%$0.070Epoch AI Benchmarking Hub2026-10-05
31Llama-3.2-3B-InstructMeta AI10.1%$0.020Epoch AI Benchmarking Hub2026-10-05
32qwen2.5-7b-instructAlibaba7.8%$0.040Epoch AI Benchmarking Hub2026-10-05
33Llama-3.2-1B-InstructMeta AI6.6%$0.020Epoch AI Benchmarking Hub2026-10-05
34Qwen2-VL-7B-InstructAlibaba3.7%$0.020Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to GPT-5.6 Luna (max) at $0.200 per million input tokens (score 45.6%).

What BALROG measures

BALROG measures long-horizon game-playing agents.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks