Terminal-Bench 2.0 leaderboard: which LLMs score highest
Terminal-Bench 2.0 (real terminal tasks (agent + model)) currently has scores for 28 models. The top score is 69.4% by Gemini 3 Pro Preview; the median model scores 37.3% and the lowest scores 9.2%.
| # | Model | Organisation | Terminal-Bench 2.0 score | Cheapest input $/M | Source | Captured |
|---|---|---|---|---|---|---|
| 1 | Gemini 3 Pro Preview | Google DeepMind | 69.4% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 2 | GPT-5.2 Codex | OpenAI | 66.5% | $1.75 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 3 | GPT-5.1 Codex Mini | OpenAI | 61.6% | $0.250 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 4 | GPT-5.1-Codex-Max | OpenAI | 60.4% | $1.25 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 5 | Claude Opus 4.5 (128k thinking) | Anthropic | 59.1% | $5.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 6 | GPT-5.1 Codex | OpenAI | 57.8% | $1.25 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 7 | grok-4-20 | xAI | 57.3% | $1.25 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 8 | GLM-5 | Z.ai (Zhipu AI) | 52.4% | $0.600 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 9 | MiniMax-M2.7 | MiniMax | 45.1% | $0.210 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 10 | Kimi K2.5 | Moonshot | 43.2% | $0.450 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 11 | MiniMax-M2.5 | MiniMax | 42.7% | $0.270 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 12 | DeepSeek-V3.2 (Thinking; Novita) | DeepSeek | 39.6% | $0.259 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 13 | claude-opus-4-1-20250805 | Anthropic | 38.0% | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 14 | Claude Opus 4.1 (unknown thinking) | Anthropic | 38.0% | $15.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 15 | MiniMax-M2.1 | MiniMax | 36.6% | $0.300 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 16 | Kimi K2 Thinking | Moonshot | 35.7% | $0.550 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 17 | GLM-4.7 | Z.ai (Zhipu AI) | 33.4% | $0.400 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 18 | MiniMax-M2 | MiniMax | 30.0% | $0.300 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 19 | claude-haiku-4-5-20251001 | Anthropic | 29.8% | $1.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 20 | Kimi K2 Instruct | Moonshot | 27.8% | $0.500 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 21 | grok-4-0709 | xAI | 27.2% | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 22 | Qwen3-Coder-480B-A35B-Instruct | Alibaba | 27.2% | $0.380 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 23 | grok-code-fast-1 | xAI | 25.8% | $0.200 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 24 | GLM-4.6 | Z.ai (Zhipu AI),Tsinghua University | 24.5% | $0.430 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 25 | Qwen3.6 35B-A3B | Alibaba | 23.0% | $0.050 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 26 | gemini-2.5-flash-preview-09-2025 | Google DeepMind | 17.1% | $0.300 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 27 | gemini-2.5-flash | Google DeepMind | 17.1% | $0.300 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 28 | Qwen3.5 9B | Alibaba | 9.2% | $0.100 | Epoch AI Benchmarking Hub | 2026-10-05 |
Compare all benchmarks side by side on the live leaderboard →
Among the ten highest scorers, the cheapest listed API price belongs to MiniMax-M2.7 at $0.210 per million input tokens (score 45.1%).
What Terminal-Bench 2.0 measures
Terminal-Bench tests agents on real tasks in a command-line environment. Scores depend on both the model and the agent harness around it.
How to read these scores
Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.
Where the data comes from
Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.