ARC-AGI-2 leaderboard: which LLMs score highest

ARC-AGI-2 (abstract reasoning, ARC Prize public eval set) currently has scores for 89 models. The top score is 95.0% by GPT-6 Astra (max); the median model scores 18.9% and the lowest scores 0.0%.

ARC-AGI-2 leaderboard: top 50 of 89 models (highest-effort row per model)
#ModelOrganisationARC-AGI-2 scoreCheapest input $/MSourceCaptured
1GPT-6 Astra (max)OpenAI95.0%$10.00ARC Prize official leaderboard2026-10-05
2GPT-6.1 Sol (max)OpenAI94.2%$2.00ARC Prize official leaderboard2026-10-05
3GPT-5.6 Sol (max)OpenAI92.5%$2.00ARC Prize official leaderboard2026-10-05
4Claude Opus 5.5 (max)Anthropic91.7%$4.00ARC Prize official leaderboard2026-10-05
5Claude Opus 5 (max)Anthropic90.4%$5.00ARC Prize official leaderboard2026-10-05
6Claude Fable 5.1 (max)Anthropic90.0%$10.00ARC Prize official leaderboard2026-10-05
7GPT-6 Sol (max)OpenAI89.6%$2.00ARC Prize official leaderboard2026-10-05
8Claude Fable 5 (max)Anthropic89.2%$10.00ARC Prize official leaderboard2026-10-05
9Gemini 3.8 Flash (high)Google DeepMind89.2%$0.750ARC Prize official leaderboard2026-10-05
10GPT-5.5 (xhigh)OpenAI85.0%$5.00ARC Prize official leaderboard2026-10-05
11Gemini 3.7 Flash (high)Google DeepMind84.6%$0.750ARC Prize official leaderboard2026-10-05
12Gemini 3 Deep Think (2/26)Google84.6%—ARC Prize official leaderboard2026-10-05
13GPT-5.5 Pro (xhigh)OpenAI84.2%$30.00ARC Prize official leaderboard2026-10-05
14GPT-5.6 Terra (max)OpenAI83.9%$2.00ARC Prize official leaderboard2026-10-05
15GPT-5.4 Pro (xhigh)OpenAI83.3%$30.00ARC Prize official leaderboard2026-10-05
16Dots3-Note Preview (Max)—76.8%—ARC Prize official leaderboard2026-10-05
17Claude 4.7 (Max)Anthropic75.8%—ARC Prize official leaderboard2026-10-05
18GPT-5.4 (xhigh)OpenAI74.0%$2.50ARC Prize official leaderboard2026-10-05
19gemini-3.5-flash-highGoogle72.1%$1.50ARC Prize official leaderboard2026-10-05
20Claude Opus 4.6 (120K, High)Anthropic69.2%$5.00ARC Prize official leaderboard2026-10-05
21Claude Opus 4.6 (120K, Max)Anthropic68.8%$5.00ARC Prize official leaderboard2026-10-05
22Grok 4.6 (xhigh)xAI67.1%$1.25ARC Prize official leaderboard2026-10-05
23Claude Opus 4.6 (120K, Medium)Anthropic66.2%$5.00ARC Prize official leaderboard2026-10-05
24GLM-5.3-Flash (max)Z.ai (Zhipu AI)65.8%$0.110ARC Prize official leaderboard2026-10-05
25Grok 4.20 (Reasoning)xAI65.1%$1.25ARC Prize official leaderboard2026-10-05
26Claude Opus 4.6 (120K, Low)Anthropic64.6%$5.00ARC Prize official leaderboard2026-10-05
27DeepSeek V4 Flash 0731 (max)DeepSeek61.4%$0.015ARC Prize official leaderboard2026-10-05
28DeepSeek V4 Pro 0813 (max)DeepSeek61.2%$0.660ARC Prize official leaderboard2026-10-05
29Gemini 3.6 Flash (High)Google60.4%$0.750ARC Prize official leaderboard2026-10-05
30Kimi K3 (max)Moonshot60.4%$1.29ARC Prize official leaderboard2026-10-05
31GPT-5.6 Luna 2026-07-30 (Max)OpenAI59.6%—ARC Prize official leaderboard2026-10-05
32GPT-5.6 Luna (max)OpenAI59.5%$0.200ARC Prize official leaderboard2026-10-05
33GPT-6 Luna (max)OpenAI59.3%$0.100ARC Prize official leaderboard2026-10-05
34Claude Sonnet 4.6 (max)Anthropic58.3%$3.00ARC Prize official leaderboard2026-10-05
35GPT-5.2 Pro (High)OpenAI54.2%$21.00ARC Prize official leaderboard2026-10-05
36GPT-5.2 (xhigh)OpenAI52.9%$1.75ARC Prize official leaderboard2026-10-05
37Grok 4.5 (high)xAI52.6%$2.00ARC Prize official leaderboard2026-10-05
38Gemini 3 Deep Think (Preview) ²Google45.1%—ARC Prize official leaderboard2026-10-05
39Qwen3.8-27B (XHigh)Alibaba42.4%$0.400ARC Prize official leaderboard2026-10-05
40Inkling Small (xhigh)Thinking Machines40.1%$0.450ARC Prize official leaderboard2026-10-05
41Opus 4.5 (Thinking, 64K)—37.6%—ARC Prize official leaderboard2026-10-05
42Gemini 3 Flash Preview (High)Google33.6%$0.500ARC Prize official leaderboard2026-10-05
43Gemini 3 ProGoogle DeepMind31.1%$2.00ARC Prize official leaderboard2026-10-05
44Opus 4.5 (Thinking, 16K)—22.8%—ARC Prize official leaderboard2026-10-05
45GPT-5.4 mini (xhigh)OpenAI18.9%$0.750ARC Prize official leaderboard2026-10-05
46GPT-5 ProOpenAI18.3%$15.00ARC Prize official leaderboard2026-10-05
47GPT-5.1 (Thinking, High)OpenAI17.6%$1.25ARC Prize official leaderboard2026-10-05
48Opus 4.5 (Thinking, 8K)—13.9%—ARC Prize official leaderboard2026-10-05
49Kimi K2.5Moonshot11.8%$0.450ARC Prize official leaderboard2026-10-05
50Gemini 3.5 Flash-Lite (High)Google10.3%$0.300ARC Prize official leaderboard2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to Gemini 3.8 Flash (high) at $0.750 per million input tokens (score 89.2%).

What ARC-AGI-2 measures

ARC-AGI-2 gives a model small coloured-grid puzzles, each needing a new rule to be inferred from a few examples. The tasks are designed to be easy for people and hard for AI, so the score reflects how well a model adapts to something it has not seen.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

ARC Prize official leaderboard. Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks