FrontierCode leaderboard: which LLMs score highest
FrontierCode (hard programming tasks) currently has scores for 15 models. The top score is 53.4% by Claude Opus 5 (max); the median model scores 31.8% and the lowest scores 8.0%.
| # | Model | Organisation | FrontierCode score | Cheapest input $/M | Source | Captured |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 (max) | Anthropic | 53.4% | $5.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 2 | GPT-6 Astra (max) | OpenAI | 53.3% | $10.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 3 | SWE-2 (max) | Cognition | 50.0% | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 4 | GPT-6 Sol (max) | OpenAI | 49.3% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 5 | GPT-6 Luna (max) | OpenAI | 42.4% | $0.100 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 6 | SWE-1.7 | Cognition | 42.0% | $0.500 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 7 | GLM-5.3 (max) | Z.ai (Zhipu AI) | 40.1% | $0.070 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 8 | GLM-5.3-Flash (max) | Z.ai (Zhipu AI) | 31.8% | $0.110 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 9 | Kimi K2.7 Code | Moonshot | 30.1% | $0.671 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 10 | Composer 2.5 | Cursor | 25.6% | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 11 | MiniMax-M3 | MiniMax | 14.7% | $0.230 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 12 | nemotron-3-ultra | Nvidia | 13.6% | $0.500 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 13 | Qwen3.7 Plus | Alibaba | 10.2% | $0.282 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 14 | SWE-1.6 | Cognition | 9.4% | $0.500 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 15 | Mistral Medium 3.5 | Mistral AI | 8.0% | $1.50 | Epoch AI Benchmarking Hub | 2026-10-05 |
Compare all benchmarks side by side on the live leaderboard →
Among the ten highest scorers, the cheapest listed API price belongs to GLM-5.3 (max) at $0.070 per million input tokens (score 40.1%).
What FrontierCode measures
FrontierCode measures hard programming tasks.
How to read these scores
Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.
Where the data comes from
Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.