DeepSWE leaderboard: which LLMs score highest

DeepSWE (long-horizon software engineering) currently has scores for 19 models. The top score is 73.8% by Gemini 3.8 Flash (high); the median model scores 67.0% and the lowest scores 30.5%.

DeepSWE leaderboard: top 19 of 19 models (highest-effort row per model)
#ModelOrganisationDeepSWE scoreCheapest input $/MSourceCaptured
1Gemini 3.8 Flash (high)Google DeepMind73.8%$0.750Epoch AI Benchmarking Hub2026-10-05
2Claude Opus 5 (max)Anthropic73.6%$5.00Epoch AI Benchmarking Hub2026-10-05
3GPT-6 Astra (max)OpenAI73.2%$10.00Epoch AI Benchmarking Hub2026-10-05
4GPT-5.6 Sol (max)OpenAI72.7%$2.00Epoch AI Benchmarking Hub2026-10-05
5Claude Fable 5 (max)Anthropic69.7%$10.00Epoch AI Benchmarking Hub2026-10-05
6GPT-5.6 Terra (max)OpenAI69.6%$2.00Epoch AI Benchmarking Hub2026-10-05
7GLM-5.3 (max)Z.ai (Zhipu AI)69.0%$0.070Epoch AI Benchmarking Hub2026-10-05
8Kimi K3 (max)Moonshot68.5%$1.29Epoch AI Benchmarking Hub2026-10-05
9GPT-5.6 Luna (max)OpenAI67.2%$0.200Epoch AI Benchmarking Hub2026-10-05
10GPT-5.5 (xhigh)OpenAI67.0%$5.00Epoch AI Benchmarking Hub2026-10-05
11Grok 4.6 (xhigh)xAI66.7%$1.25Epoch AI Benchmarking Hub2026-10-05
12Gemini 3.7 Flash (high)Google DeepMind65.3%$0.750Epoch AI Benchmarking Hub2026-10-05
13GLM-5.3-Flash (max)Z.ai (Zhipu AI)63.4%$0.110Epoch AI Benchmarking Hub2026-10-05
14Qwen3.8 Max (xhigh)Alibaba57.5%$1.65Epoch AI Benchmarking Hub2026-10-05
15Muse Spark 1.2 (xhigh)Meta AI54.9%$1.25Epoch AI Benchmarking Hub2026-10-05
16Claude Sonnet 5 (max)Anthropic53.8%$2.00Epoch AI Benchmarking Hub2026-10-05
17Grok 4.5 (high)xAI53.8%$2.00Epoch AI Benchmarking Hub2026-10-05
18GPT-5.4 (xhigh)OpenAI51.8%$2.50Epoch AI Benchmarking Hub2026-10-05
19Kimi K2.7 CodeMoonshot30.5%$0.671Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to GLM-5.3 (max) at $0.070 per million input tokens (score 69.0%).

What DeepSWE measures

DeepSWE measures long-horizon software engineering.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks