Best LLMs for Coding: Benchmark Ranking (2026)

GPT-6 Astra (max) leads this ranking with a score of 60.7 (50 is the neutral midpoint). This page ranks models on software-engineering benchmarks: resolving real GitHub issues (SWE-bench family), editing code with a tool (Aider), terminal and agentic coding tasks, competitive programming (ALE-Bench) and crowd-voted coding answers (LMArena Code). Data last captured 2026-10-05.

Best LLMs for Coding: top 30
#ModelOrganisationScoreOverall rankCheapest input $/M
1GPT-6 Astra (max)OpenAI60.7#2$10.00
2Claude Opus 5.5 (max)Anthropic60.3#1$4.00
3GPT-5.5 (xhigh)OpenAI57.5#10$5.00
4Claude Fable 5.1 (max)Anthropic57.4#3$10.00
5GPT-6 Sol (max)OpenAI57.1#8$2.00
6Claude Sonnet 5.5 (max)Anthropic55.7#4$2.00
7Claude Opus 5 (max)Anthropic55.5#6$5.00
8GPT-6.1 Sol (max)OpenAI55.4#5$2.00
9GPT-5.2 CodexOpenAI54.8#77$1.75
10Gemini 3 Pro PreviewGoogle DeepMind54.1#30$2.00
11GPT-5.4 (xhigh)OpenAI54.1#13$2.50
12GPT-5.6 Sol (max)OpenAI53.8#7$2.00
13Claude Fable 5 (max)Anthropic53.5#9$10.00
14GPT-5.6 Terra (max)OpenAI53.4#11$2.00
15Kimi K3 (max)Moonshot53.4#17$1.29
16Claude 3.5-Sonnet-20241022Anthropic52.9#24$3.00
17Muse Spark 1.3 (max)Meta AI52.8#20$1.25
18Kimi K2.6Moonshot52.5#351$0.650
19GLM-5.3 (max)Z.ai (Zhipu AI)52.3#31$0.070
20GPT-5.1 CodexOpenAI51.6#174$1.25
21GPT-5.6 Luna (max)OpenAI51.6#19$0.200
22GPT-5.1 (high)OpenAI51.5#52$1.25
23GLM-5Z.ai (Zhipu AI)51.4#121$0.600
24deepseek-v4.1-flash-maxDeepSeek51.3#128—
25GPT-6 Luna (max)OpenAI51.3#34$0.100
26grok-4-20xAI51.3#219$1.25
27Qwen3.7 MaxAlibaba51.1#69$1.25
28GPT-5 (high)OpenAI51.0#147$1.25
29Grok 4.7 (xhigh)xAI50.8#73$2.00
30claude-3-7-sonnet-20250219Anthropic50.8#382—

How this ranking is built

This ranking uses only software-engineering benchmarks: resolving real GitHub issues (SWE-bench family), editing code with a tool (Aider), terminal and agentic coding tasks, competitive programming (ALE-Bench) and crowd-voted coding answers (LMArena Code). For each model we take its z-score on every one of these benchmarks it has been measured on and shrink the average toward 50 when it has few results; models need at least three of them to appear. Benchmarks used: SWE-bench Verified, SWE-bench Lite, Aider Polyglot, Aider Code Editing, Terminal-Bench 2.0, Terminal-Bench 4.0, DeepSWE, FrontierSWE, ALE-Bench, FrontierCode, LiveBench Coding, LiveBench Agentic Coding, LMArena Code Elo, GSO-Bench, WebDev Arena and SciCode.

More rankings