SWE-bench Verified leaderboard: which LLMs score highest

SWE-bench Verified (real GitHub issues resolved) currently has scores for 73 models. The top score is 83.5% by Claude Opus 4.7 (max); the median model scores 56.4% and the lowest scores 0.4%.

SWE-bench Verified leaderboard: top 50 of 73 models (highest-effort row per model)
#ModelOrganisationSWE-bench Verified scoreCheapest input $/MSourceCaptured
1Claude Opus 4.7 (max)Anthropic83.5%$5.00Epoch AI Benchmarking Hub2026-10-05
2DeepSeek v4 Pro (max)DeepSeek77.6%$0.435Epoch AI Benchmarking Hub2026-10-05
3Qwen3.7 MaxAlibaba77.3%$1.25Epoch AI Benchmarking Hub2026-10-05
4Kimi K2.6Moonshot76.7%$0.650Epoch AI Benchmarking Hub2026-10-05
5Qwen 3.6 Max (Preview)Alibaba76.7%—Epoch AI Benchmarking Hub2026-10-05
6Gemini 3 FlashGoogle DeepMind75.8%$0.500SWE-bench Verified leaderboard2026-10-05
7MiniMax-M2.5MiniMax75.8%$0.270SWE-bench Verified leaderboard2026-10-05
8Claude 4.6 OpusAnthropic75.6%—SWE-bench Verified leaderboard2026-10-05
9Claude 4.5 OpusAnthropic74.4%—SWE-bench Verified leaderboard2026-10-05
10Gemini 3 Pro PreviewGoogle DeepMind74.2%$2.00SWE-bench Verified leaderboard2026-10-05
11GLM-5.1Z.ai (Zhipu AI)74.2%$1.05Epoch AI Benchmarking Hub2026-10-05
12GPT-5 (high)OpenAI73.6%$1.25Epoch AI Benchmarking Hub2026-10-05
13claude-opus-4-1-20250805Anthropic73.3%—Epoch AI Benchmarking Hub2026-10-05
14GLM-5Z.ai (Zhipu AI)72.8%$0.600SWE-bench Verified leaderboard2026-10-05
15GPT-5.2 CodexOpenAI72.8%$1.75SWE-bench Verified leaderboard2026-10-05
16Claude 4 SonnetAnthropic70.8%$3.00SWE-bench Verified leaderboard2026-10-05
17Kimi K2.5Moonshot70.8%$0.450SWE-bench Verified leaderboard2026-10-05
18Claude Opus 4Anthropic70.7%$15.00Epoch AI Benchmarking Hub2026-10-05
19Claude 4.5 SonnetAnthropic70.6%$3.00SWE-bench Verified leaderboard2026-10-05
20DeepSeek-V3.2DeepSeek70.0%$0.259SWE-bench Verified leaderboard2026-10-05
21Gemini 3 ProGoogle DeepMind69.6%$2.00SWE-bench Verified leaderboard2026-10-05
22GPT-5.1 (high)OpenAI68.0%$1.25Epoch AI Benchmarking Hub2026-10-05
23Claude 4 OpusAnthropic67.6%$5.00SWE-bench Verified leaderboard2026-10-05
24Claude 4.5 HaikuAnthropic66.6%$1.00SWE-bench Verified leaderboard2026-10-05
25GPT-5.1 CodexOpenAI66.0%$1.25SWE-bench Verified leaderboard2026-10-05
26Kimi K2Moonshot AI65.4%$0.550SWE-bench Verified leaderboard2026-10-05
27o1-previewOpenAI64.6%$15.00SWE-bench Verified leaderboard2026-10-05
28Kimi K2 ThinkingMoonshot63.4%$0.550SWE-bench Verified leaderboard2026-10-05
29Claude 3.7 Sonnet w/ Review HeavyAnthropic62.4%—SWE-bench Verified leaderboard2026-10-05
30swe-searchAnthropic62.2%—SWE-bench Verified leaderboard2026-10-05
31MiniMax-M2MiniMax61.0%$0.300SWE-bench Verified leaderboard2026-10-05
32claude-3-7-sonnet-20250219Anthropic61.0%—Epoch AI Benchmarking Hub2026-10-05
33Qwen3-Coder-30B-A3B-InstructQwen60.4%$0.070SWE-bench Verified leaderboard2026-10-05
34DeepSeek V3.2 Reasonerdeepseek60.0%—SWE-bench Verified leaderboard2026-10-05
35TTS(Bo16)OpenAI58.8%—SWE-bench Verified leaderboard2026-10-05
36Qwen 3.6 Plus (2026-04-02)Alibaba57.9%—Epoch AI Benchmarking Hub2026-10-05
37Devstral Small (2512)mistral56.4%$0.070SWE-bench Verified leaderboard2026-10-05
38GLM-4.6Z.ai (Zhipu AI),Tsinghua University55.4%$0.430SWE-bench Verified leaderboard2026-10-05
39Qwen3-Coder-480B-A35B-InstructAlibaba55.4%$0.380SWE-bench Verified leaderboard2026-10-05
40glm-4.5Z.ai (Zhipu AI),Tsinghua University54.2%$0.400SWE-bench Verified leaderboard2026-10-05
41Devstral (2512)mistral53.8%$0.400SWE-bench Verified leaderboard2026-10-05
42Frogboss 32B 2510—53.6%—SWE-bench Verified leaderboard2026-10-05
43Claude 3.5-Sonnet-20241022Anthropic51.6%$3.00SWE-bench Verified leaderboard2026-10-05
44TTS(Bo8)Qwen47.0%—SWE-bench Verified leaderboard2026-10-05
45DevStral Small 2505Mistral46.8%—SWE-bench Verified leaderboard2026-10-05
46Frogmini 14B 2510—45.0%—SWE-bench Verified leaderboard2026-10-05
47Gemini 2.0 Flash (Experimental)Google DeepMind44.2%$0.150SWE-bench Verified leaderboard2026-10-05
48Kimi K2 InstructMoonshot43.8%$0.500SWE-bench Verified leaderboard2026-10-05
49Amazon.nova Premier v1:0—42.4%—SWE-bench Verified leaderboard2026-10-05
50DeepSWE-Preview—42.2%—SWE-bench Verified leaderboard2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to MiniMax-M2.5 at $0.270 per million input tokens (score 75.8%).

What SWE-bench Verified measures

SWE-bench Verified asks a model, wrapped in an agent scaffold, to fix real issues from open-source Python projects. A fix counts only if the project's own tests pass. The Verified subset was hand-checked by people to make sure each task is solvable.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0), SWE-bench Verified leaderboard (unverified) and SWE-bench Verified leaderboard. Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks