SWE-bench Lite leaderboard: which LLMs score highest

SWE-bench Lite (real GitHub issues resolved) currently has scores for 20 models. The top score is 56.7% by Claude 4 Sonnet; the median model scores 24.7% and the lowest scores 0.3%.

SWE-bench Lite leaderboard: top 20 of 20 models (highest-effort row per model)
#ModelOrganisationSWE-bench Lite scoreCheapest input $/MSourceCaptured
1Claude 4 SonnetAnthropic56.7%$3.00SWE-bench Lite leaderboard2026-10-05
2Qwen3-Coder-30B-A3B-InstructQwen49.7%$0.070SWE-bench Lite leaderboard2026-10-05
3CodeAct v2.1 (claude-3-5-sonnet-20241022)Anthropic41.7%—SWE-bench Lite leaderboard2026-10-05
4GPT-4 (0806)OpenAI39.7%$30.00SWE-bench Lite leaderboard2026-10-05
5Claude 3.5 SonnetAnthropic39.0%$3.00SWE-bench Lite leaderboard2026-10-05
6DeepSeek-V3DeepSeek30.7%$0.200SWE-bench Lite leaderboard2026-10-05
7o3-mini_1.0OpenAI30.3%—SWE-bench Lite leaderboard2026-10-05
8Claude 3.5-Sonnet-20241022Anthropic30.0%$3.00SWE-bench Lite leaderboard2026-10-05
9CodeAct v1.8Anthropic26.7%—SWE-bench Lite leaderboard2026-10-05
10GPT-4oOpenAI24.7%$2.50SWE-bench Lite leaderboard2026-10-05
11Qwen2.5 (7B + 72B)Qwen24.7%—SWE-bench Lite leaderboard2026-10-05
12GPT-4 (0613)OpenAI23.7%$30.00SWE-bench Lite leaderboard2026-10-05
13GPT-4 (0125)OpenAI19.0%$30.00SWE-bench Lite leaderboard2026-10-05
14GPT-4 (1106)OpenAI18.0%$30.00SWE-bench Lite leaderboard2026-10-05
15MCTS Refine 7B—16.3%—SWE-bench Lite leaderboard2026-10-05
16Claude 3 OpusAnthropic11.7%$15.00SWE-bench Lite leaderboard2026-10-05
17Claude 2Anthropic3.0%—SWE-bench Lite leaderboard2026-10-05
18SWE-Llama 7BMeta1.3%—SWE-bench Lite leaderboard2026-10-05
19SWE-Llama 13B—1.0%—SWE-bench Lite leaderboard2026-10-05
20GPT-3.5OpenAI0.3%—SWE-bench Lite leaderboard2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to Qwen3-Coder-30B-A3B-Instruct at $0.070 per million input tokens (score 49.7%).

What SWE-bench Lite measures

SWE-bench Lite is a smaller, cheaper subset of SWE-bench with 300 tasks, built so that more teams can run it.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

SWE-bench Lite leaderboard and SWE-bench Lite leaderboard (unverified). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks