Terminal-Bench 4.0 leaderboard: which LLMs score highest
Terminal-Bench 4.0 (real terminal tasks, agent+model config) currently has scores for 15 models. The top score is 58.2% by Codex + GPT-6 Astra; the median model scores 23.6% and the lowest scores 11.2%.
| # | Model | Organisation | Terminal-Bench 4.0 score | Cheapest input $/M | Source | Captured |
|---|---|---|---|---|---|---|
| 1 | Codex + GPT-6 Astra | โ | 58.2% | โ | Terminal-Bench 4.0 official leaderboard | 2026-10-05 |
| 2 | Claude Code + Fable 5.1 | โ | 57.9% | โ | Terminal-Bench 4.0 official leaderboard | 2026-10-05 |
| 3 | Claude Code + Opus 5 | โ | 53.9% | โ | Terminal-Bench 4.0 official leaderboard | 2026-10-05 |
| 4 | Claude Code + Fable 5 | โ | 44.5% | โ | Terminal-Bench 4.0 official leaderboard | 2026-10-05 |
| 5 | Claude Code + GLM-5.3 | โ | 41.8% | โ | Terminal-Bench 4.0 official leaderboard | 2026-10-05 |
| 6 | Grok Build + Grok 4.7 | โ | 37.6% | โ | Terminal-Bench 4.0 official leaderboard | 2026-10-05 |
| 7 | Codex + GPT-5.6 Sol | โ | 37.3% | โ | Terminal-Bench 4.0 official leaderboard | 2026-10-05 |
| 8 | Claude Code + Opus 4.8 | โ | 23.6% | โ | Terminal-Bench 4.0 official leaderboard | 2026-10-05 |
| 9 | Codex + GPT-5.6 Terra | โ | 21.5% | โ | Terminal-Bench 4.0 official leaderboard | 2026-10-05 |
| 10 | Grok Build + Grok 4.6 | โ | 20.3% | โ | Terminal-Bench 4.0 official leaderboard | 2026-10-05 |
| 11 | mini-SWE-agent + Gemini 3.8 Flash | โ | 19.1% | โ | Terminal-Bench 4.0 official leaderboard | 2026-10-05 |
| 12 | Codex + GPT-5.6 Luna | โ | 17.3% | โ | Terminal-Bench 4.0 official leaderboard | 2026-10-05 |
| 13 | Claude Code + Sonnet 5 | โ | 12.4% | โ | Terminal-Bench 4.0 official leaderboard | 2026-10-05 |
| 14 | Grok Build + Grok 4.5 | โ | 12.4% | โ | Terminal-Bench 4.0 official leaderboard | 2026-10-05 |
| 15 | mini-SWE-agent + Gemini 3.7 Flash | โ | 11.2% | โ | Terminal-Bench 4.0 official leaderboard | 2026-10-05 |
Compare all benchmarks side by side on the live leaderboard โ
What Terminal-Bench 4.0 measures
Terminal-Bench tests agents on real tasks in a command-line environment. Scores depend on both the model and the agent harness around it.
How to read these scores
Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.
Where the data comes from
Terminal-Bench 4.0 official leaderboard. Scores were last captured 2026-10-05; the Captured column gives each row's date.