LiveBench Agentic Coding leaderboard: which LLMs score highest
LiveBench Agentic Coding (LiveBench release 2026_06_25) currently has scores for 17 models. The top score is 56.1% by claude-opus-4-8-xhigh-effort; the median model scores 45.8% and the lowest scores 39.3%.
| # | Model | Organisation | LiveBench Agentic Coding score | Cheapest input $/M | Source | Captured |
|---|---|---|---|---|---|---|
| 1 | claude-opus-4-8-xhigh-effort | Anthropic | 56.1% | $5.00 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 2 | GPT-5.4 (xhigh) | OpenAI | 53.8% | $2.50 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 3 | GPT-5.5 (xhigh) | OpenAI | 52.1% | $5.00 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 4 | gpt-5.2-2025-12-11-high | OpenAI | 50.3% | $1.75 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 5 | GPT-5.2 Codex | OpenAI | 49.4% | $1.75 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 6 | gemini-3.5-flash-high | 49.0% | $1.50 | LiveBench 2026_06_25 official CSV | 2026-10-05 | |
| 7 | kimi-k2.6-thinking | Moonshot AI | 46.9% | $0.650 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 8 | GPT-5.4 nano (xhigh) | OpenAI | 46.8% | $0.200 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 9 | grok-build-0.1 | xAI | 45.8% | $1.00 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 10 | Kimi K2.7 Code | Moonshot | 45.7% | $0.671 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 11 | gemini-3.1-pro-preview-high | 45.4% | $2.00 | LiveBench 2026_06_25 official CSV | 2026-10-05 | |
| 12 | Qwen3.7 Max | Alibaba | 43.6% | $1.25 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 13 | GPT-5.4 mini (xhigh) | OpenAI | 41.7% | $0.750 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 14 | qwen3.6-plus | Alibaba | 41.4% | $0.325 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 15 | MiniMax-M3 | MiniMax | 40.7% | $0.230 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 16 | claude-opus-4-5-20251101-thinking-64k-high-effort | Anthropic | 39.7% | $5.00 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 17 | Qwen3.6 27B | Alibaba | 39.3% | $0.150 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
Compare all benchmarks side by side on the live leaderboard →
Among the ten highest scorers, the cheapest listed API price belongs to GPT-5.4 nano (xhigh) at $0.200 per million input tokens (score 46.8%).
What LiveBench Agentic Coding measures
LiveBench Agentic Coding measures LiveBench release 2026_06_25.
How to read these scores
Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.
Where the data comes from
LiveBench 2026_06_25 official CSV. Scores were last captured 2026-10-05; the Captured column gives each row's date.