Best LLMs for Coding: Benchmark Ranking (2026)
GPT-6 Astra (max) leads this ranking with a score of 60.7 (50 is the neutral midpoint). This page ranks models on software-engineering benchmarks: resolving real GitHub issues (SWE-bench family), editing code with a tool (Aider), terminal and agentic coding tasks, competitive programming (ALE-Bench) and crowd-voted coding answers (LMArena Code). Data last captured 2026-10-05.
| # | Model | Organisation | Score | Overall rank | Cheapest input $/M |
|---|---|---|---|---|---|
| 1 | GPT-6 Astra (max) | OpenAI | 60.7 | #2 | $10.00 |
| 2 | Claude Opus 5.5 (max) | Anthropic | 60.3 | #1 | $4.00 |
| 3 | GPT-5.5 (xhigh) | OpenAI | 57.5 | #10 | $5.00 |
| 4 | Claude Fable 5.1 (max) | Anthropic | 57.4 | #3 | $10.00 |
| 5 | GPT-6 Sol (max) | OpenAI | 57.1 | #8 | $2.00 |
| 6 | Claude Sonnet 5.5 (max) | Anthropic | 55.7 | #4 | $2.00 |
| 7 | Claude Opus 5 (max) | Anthropic | 55.5 | #6 | $5.00 |
| 8 | GPT-6.1 Sol (max) | OpenAI | 55.4 | #5 | $2.00 |
| 9 | GPT-5.2 Codex | OpenAI | 54.8 | #77 | $1.75 |
| 10 | Gemini 3 Pro Preview | Google DeepMind | 54.1 | #30 | $2.00 |
| 11 | GPT-5.4 (xhigh) | OpenAI | 54.1 | #13 | $2.50 |
| 12 | GPT-5.6 Sol (max) | OpenAI | 53.8 | #7 | $2.00 |
| 13 | Claude Fable 5 (max) | Anthropic | 53.5 | #9 | $10.00 |
| 14 | GPT-5.6 Terra (max) | OpenAI | 53.4 | #11 | $2.00 |
| 15 | Kimi K3 (max) | Moonshot | 53.4 | #17 | $1.29 |
| 16 | Claude 3.5-Sonnet-20241022 | Anthropic | 52.9 | #24 | $3.00 |
| 17 | Muse Spark 1.3 (max) | Meta AI | 52.8 | #20 | $1.25 |
| 18 | Kimi K2.6 | Moonshot | 52.5 | #351 | $0.650 |
| 19 | GLM-5.3 (max) | Z.ai (Zhipu AI) | 52.3 | #31 | $0.070 |
| 20 | GPT-5.1 Codex | OpenAI | 51.6 | #174 | $1.25 |
| 21 | GPT-5.6 Luna (max) | OpenAI | 51.6 | #19 | $0.200 |
| 22 | GPT-5.1 (high) | OpenAI | 51.5 | #52 | $1.25 |
| 23 | GLM-5 | Z.ai (Zhipu AI) | 51.4 | #121 | $0.600 |
| 24 | deepseek-v4.1-flash-max | DeepSeek | 51.3 | #128 | — |
| 25 | GPT-6 Luna (max) | OpenAI | 51.3 | #34 | $0.100 |
| 26 | grok-4-20 | xAI | 51.3 | #219 | $1.25 |
| 27 | Qwen3.7 Max | Alibaba | 51.1 | #69 | $1.25 |
| 28 | GPT-5 (high) | OpenAI | 51.0 | #147 | $1.25 |
| 29 | Grok 4.7 (xhigh) | xAI | 50.8 | #73 | $2.00 |
| 30 | claude-3-7-sonnet-20250219 | Anthropic | 50.8 | #382 | — |
How this ranking is built
This ranking uses only software-engineering benchmarks: resolving real GitHub issues (SWE-bench family), editing code with a tool (Aider), terminal and agentic coding tasks, competitive programming (ALE-Bench) and crowd-voted coding answers (LMArena Code). For each model we take its z-score on every one of these benchmarks it has been measured on and shrink the average toward 50 when it has few results; models need at least three of them to appear. Benchmarks used: SWE-bench Verified, SWE-bench Lite, Aider Polyglot, Aider Code Editing, Terminal-Bench 2.0, Terminal-Bench 4.0, DeepSWE, FrontierSWE, ALE-Bench, FrontierCode, LiveBench Coding, LiveBench Agentic Coding, LMArena Code Elo, GSO-Bench, WebDev Arena and SciCode.