Terminal-Bench 2.0 leaderboard: which LLMs score highest

Terminal-Bench 2.0 (real terminal tasks (agent + model)) currently has scores for 28 models. The top score is 69.4% by Gemini 3 Pro Preview; the median model scores 37.3% and the lowest scores 9.2%.

Terminal-Bench 2.0 leaderboard: top 28 of 28 models (highest-effort row per model)
#ModelOrganisationTerminal-Bench 2.0 scoreCheapest input $/MSourceCaptured
1Gemini 3 Pro PreviewGoogle DeepMind69.4%$2.00Epoch AI Benchmarking Hub2026-10-05
2GPT-5.2 CodexOpenAI66.5%$1.75Epoch AI Benchmarking Hub2026-10-05
3GPT-5.1 Codex MiniOpenAI61.6%$0.250Epoch AI Benchmarking Hub2026-10-05
4GPT-5.1-Codex-MaxOpenAI60.4%$1.25Epoch AI Benchmarking Hub2026-10-05
5Claude Opus 4.5 (128k thinking)Anthropic59.1%$5.00Epoch AI Benchmarking Hub2026-10-05
6GPT-5.1 CodexOpenAI57.8%$1.25Epoch AI Benchmarking Hub2026-10-05
7grok-4-20xAI57.3%$1.25Epoch AI Benchmarking Hub2026-10-05
8GLM-5Z.ai (Zhipu AI)52.4%$0.600Epoch AI Benchmarking Hub2026-10-05
9MiniMax-M2.7MiniMax45.1%$0.210Epoch AI Benchmarking Hub2026-10-05
10Kimi K2.5Moonshot43.2%$0.450Epoch AI Benchmarking Hub2026-10-05
11MiniMax-M2.5MiniMax42.7%$0.270Epoch AI Benchmarking Hub2026-10-05
12DeepSeek-V3.2 (Thinking; Novita)DeepSeek39.6%$0.259Epoch AI Benchmarking Hub2026-10-05
13claude-opus-4-1-20250805Anthropic38.0%—Epoch AI Benchmarking Hub2026-10-05
14Claude Opus 4.1 (unknown thinking)Anthropic38.0%$15.00Epoch AI Benchmarking Hub2026-10-05
15MiniMax-M2.1MiniMax36.6%$0.300Epoch AI Benchmarking Hub2026-10-05
16Kimi K2 ThinkingMoonshot35.7%$0.550Epoch AI Benchmarking Hub2026-10-05
17GLM-4.7Z.ai (Zhipu AI)33.4%$0.400Epoch AI Benchmarking Hub2026-10-05
18MiniMax-M2MiniMax30.0%$0.300Epoch AI Benchmarking Hub2026-10-05
19claude-haiku-4-5-20251001Anthropic29.8%$1.00Epoch AI Benchmarking Hub2026-10-05
20Kimi K2 InstructMoonshot27.8%$0.500Epoch AI Benchmarking Hub2026-10-05
21grok-4-0709xAI27.2%—Epoch AI Benchmarking Hub2026-10-05
22Qwen3-Coder-480B-A35B-InstructAlibaba27.2%$0.380Epoch AI Benchmarking Hub2026-10-05
23grok-code-fast-1xAI25.8%$0.200Epoch AI Benchmarking Hub2026-10-05
24GLM-4.6Z.ai (Zhipu AI),Tsinghua University24.5%$0.430Epoch AI Benchmarking Hub2026-10-05
25Qwen3.6 35B-A3BAlibaba23.0%$0.050Epoch AI Benchmarking Hub2026-10-05
26gemini-2.5-flash-preview-09-2025Google DeepMind17.1%$0.300Epoch AI Benchmarking Hub2026-10-05
27gemini-2.5-flashGoogle DeepMind17.1%$0.300Epoch AI Benchmarking Hub2026-10-05
28Qwen3.5 9BAlibaba9.2%$0.100Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to MiniMax-M2.7 at $0.210 per million input tokens (score 45.1%).

What Terminal-Bench 2.0 measures

Terminal-Bench tests agents on real tasks in a command-line environment. Scores depend on both the model and the agent harness around it.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks