SWE-bench Lite leaderboard: which LLMs score highest
SWE-bench Lite (real GitHub issues resolved) currently has scores for 20 models. The top score is 56.7% by Claude 4 Sonnet; the median model scores 24.7% and the lowest scores 0.3%.
| # | Model | Organisation | SWE-bench Lite score | Cheapest input $/M | Source | Captured |
|---|---|---|---|---|---|---|
| 1 | Claude 4 Sonnet | Anthropic | 56.7% | $3.00 | SWE-bench Lite leaderboard | 2026-10-05 |
| 2 | Qwen3-Coder-30B-A3B-Instruct | Qwen | 49.7% | $0.070 | SWE-bench Lite leaderboard | 2026-10-05 |
| 3 | CodeAct v2.1 (claude-3-5-sonnet-20241022) | Anthropic | 41.7% | — | SWE-bench Lite leaderboard | 2026-10-05 |
| 4 | GPT-4 (0806) | OpenAI | 39.7% | $30.00 | SWE-bench Lite leaderboard | 2026-10-05 |
| 5 | Claude 3.5 Sonnet | Anthropic | 39.0% | $3.00 | SWE-bench Lite leaderboard | 2026-10-05 |
| 6 | DeepSeek-V3 | DeepSeek | 30.7% | $0.200 | SWE-bench Lite leaderboard | 2026-10-05 |
| 7 | o3-mini_1.0 | OpenAI | 30.3% | — | SWE-bench Lite leaderboard | 2026-10-05 |
| 8 | Claude 3.5-Sonnet-20241022 | Anthropic | 30.0% | $3.00 | SWE-bench Lite leaderboard | 2026-10-05 |
| 9 | CodeAct v1.8 | Anthropic | 26.7% | — | SWE-bench Lite leaderboard | 2026-10-05 |
| 10 | GPT-4o | OpenAI | 24.7% | $2.50 | SWE-bench Lite leaderboard | 2026-10-05 |
| 11 | Qwen2.5 (7B + 72B) | Qwen | 24.7% | — | SWE-bench Lite leaderboard | 2026-10-05 |
| 12 | GPT-4 (0613) | OpenAI | 23.7% | $30.00 | SWE-bench Lite leaderboard | 2026-10-05 |
| 13 | GPT-4 (0125) | OpenAI | 19.0% | $30.00 | SWE-bench Lite leaderboard | 2026-10-05 |
| 14 | GPT-4 (1106) | OpenAI | 18.0% | $30.00 | SWE-bench Lite leaderboard | 2026-10-05 |
| 15 | MCTS Refine 7B | — | 16.3% | — | SWE-bench Lite leaderboard | 2026-10-05 |
| 16 | Claude 3 Opus | Anthropic | 11.7% | $15.00 | SWE-bench Lite leaderboard | 2026-10-05 |
| 17 | Claude 2 | Anthropic | 3.0% | — | SWE-bench Lite leaderboard | 2026-10-05 |
| 18 | SWE-Llama 7B | Meta | 1.3% | — | SWE-bench Lite leaderboard | 2026-10-05 |
| 19 | SWE-Llama 13B | — | 1.0% | — | SWE-bench Lite leaderboard | 2026-10-05 |
| 20 | GPT-3.5 | OpenAI | 0.3% | — | SWE-bench Lite leaderboard | 2026-10-05 |
Compare all benchmarks side by side on the live leaderboard →
Among the ten highest scorers, the cheapest listed API price belongs to Qwen3-Coder-30B-A3B-Instruct at $0.070 per million input tokens (score 49.7%).
What SWE-bench Lite measures
SWE-bench Lite is a smaller, cheaper subset of SWE-bench with 300 tasks, built so that more teams can run it.
How to read these scores
Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.
Where the data comes from
SWE-bench Lite leaderboard and SWE-bench Lite leaderboard (unverified). Scores were last captured 2026-10-05; the Captured column gives each row's date.