LiveBench Coding leaderboard: which LLMs score highest
LiveBench Coding (LiveBench release 2026_06_25) currently has scores for 17 models. The top score is 83.6% by GPT-5.2 Codex; the median model scores 76.5% and the lowest scores 65.4%.
| # | Model | Organisation | LiveBench Coding score | Cheapest input $/M | Source | Captured |
|---|---|---|---|---|---|---|
| 1 | GPT-5.2 Codex | OpenAI | 83.6% | $1.75 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 2 | GPT-5.5 (xhigh) | OpenAI | 82.1% | $5.00 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 3 | claude-opus-4-5-20251101-thinking-64k-high-effort | Anthropic | 79.7% | $5.00 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 4 | claude-opus-4-8-xhigh-effort | Anthropic | 79.3% | $5.00 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 5 | kimi-k2.6-thinking | Moonshot AI | 78.6% | $0.650 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 6 | gemini-3.5-flash-high | 78.2% | $1.50 | LiveBench 2026_06_25 official CSV | 2026-10-05 | |
| 7 | qwen3.6-plus | Alibaba | 78.2% | $0.325 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 8 | GPT-5.4 (xhigh) | OpenAI | 77.5% | $2.50 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 9 | gemini-3.1-pro-preview-high | 76.5% | $2.00 | LiveBench 2026_06_25 official CSV | 2026-10-05 | |
| 10 | gpt-5.2-2025-12-11-high | OpenAI | 76.1% | $1.75 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 11 | Qwen3.7 Max | Alibaba | 74.2% | $1.25 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 12 | Kimi K2.7 Code | Moonshot | 74.0% | $0.671 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 13 | Qwen3.6 27B | Alibaba | 71.8% | $0.150 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 14 | GPT-5.4 mini (xhigh) | OpenAI | 71.6% | $0.750 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 15 | GPT-5.4 nano (xhigh) | OpenAI | 70.8% | $0.200 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 16 | MiniMax-M3 | MiniMax | 68.2% | $0.230 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 17 | grok-build-0.1 | xAI | 65.4% | $1.00 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
Compare all benchmarks side by side on the live leaderboard →
Among the ten highest scorers, the cheapest listed API price belongs to qwen3.6-plus at $0.325 per million input tokens (score 78.2%).
What LiveBench Coding measures
LiveBench refreshes its questions regularly from recent material to limit training-data contamination, and scores answers automatically against ground truth. This is its coding category.
How to read these scores
Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.
Where the data comes from
LiveBench 2026_06_25 official CSV. Scores were last captured 2026-10-05; the Captured column gives each row's date.