METR time horizon leaderboard: which LLMs score highest
METR time horizon (length of task (minutes) done at 50% success) currently has scores for 34 models. The top score is 1,045 min by Claude Mythos Preview (Early); the median model scores 20 min and the lowest scores 0 min.
| # | Model | Organisation | METR time horizon score | Cheapest input $/M | Source | Captured |
|---|---|---|---|---|---|---|
| 1 | Claude Mythos Preview (Early) | Anthropic | 1,045 min | $10.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 2 | GPT-5.4 (xhigh) | OpenAI | 342 min | $2.50 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 3 | Gemini 3 Pro Preview | Google DeepMind | 224 min | $2.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 4 | GPT-5.1-Codex-Max | OpenAI | 224 min | $1.25 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 5 | GPT-5 (high) | OpenAI | 203 min | $1.25 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 6 | claude-opus-4-1-20250805_16K | Anthropic | 114 min | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 7 | grok-4-0709 | xAI | 110 min | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 8 | claude-opus-4-1-20250805 | Anthropic | 100 min | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 9 | Claude Opus 4 | Anthropic | 100 min | $15.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 10 | claude-sonnet-4-20250514_16K | Anthropic | 75 min | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 11 | claude-3-7-sonnet-20250219 | Anthropic | 60 min | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 12 | Kimi K2 Thinking | Moonshot | 54 min | $0.550 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 13 | Gemini 2.5 Pro Preview (Jun 2025) | Google DeepMind | 39 min | $1.25 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 14 | DeepSeek-R1 (May 2025) | DeepSeek | 31 min | $0.400 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 15 | DeepSeek-R1 | DeepSeek | 27 min | $0.400 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 16 | DeepSeek-V3 (Mar 2025) | DeepSeek | 23 min | $0.200 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 17 | Claude 3.5 Sonnet (Oct 2024) | Anthropic | 21 min | $3.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 18 | o1-preview | OpenAI | 20 min | $15.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 19 | DeepSeek-V3 | DeepSeek | 18 min | $0.200 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 20 | Claude 3.5 Sonnet (Jun 2024) | Anthropic | 11 min | $3.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 21 | GPT-4o (Nov 2024) | OpenAI | 9 min | $2.50 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 22 | GPT-4o (Aug 2024) | OpenAI | 7 min | $2.50 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 23 | gpt-4-turbo-2024-04-09 | OpenAI | 7 min | $10.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 24 | GPT-4 Turbo Preview (January 2024) | OpenAI | 5 min | $10.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 25 | GPT-4 (Mar 2023) | OpenAI | 5 min | $30.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 26 | qwen2.5-72b-instruct | Alibaba | 5 min | $0.120 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 27 | GPT-4 Turbo Preview (Nov 2023) | OpenAI | 4 min | $10.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 28 | GPT-4 (Jun 2023) | OpenAI | 4 min | $30.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 29 | claude-3-opus-20240229 | Anthropic | 4 min | $15.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 30 | gpt-4-turbo | OpenAI | 4 min | $10.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 31 | qwen2-72b-instruct | Alibaba | 2 min | $0.900 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 32 | gpt-3.5-turbo-instruct | OpenAI | 1 min | $1.50 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 33 | davinci-002 | OpenAI | 0 min | $2.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 34 | gpt2-xl | OpenAI | 0 min | — | Epoch AI Benchmarking Hub | 2026-10-05 |
Compare all benchmarks side by side on the live leaderboard →
Among the ten highest scorers, the cheapest listed API price belongs to GPT-5.1-Codex-Max at $1.25 per million input tokens (score 224 min).
What METR time horizon measures
METR time horizon measures length of task (minutes) done at 50% success.
How to read these scores
Higher scores are better; the scale is specific to this benchmark, so compare models only on this board. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.
Where the data comes from
Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.