METR time horizon leaderboard: which LLMs score highest

METR time horizon (length of task (minutes) done at 50% success) currently has scores for 34 models. The top score is 1,045 min by Claude Mythos Preview (Early); the median model scores 20 min and the lowest scores 0 min.

METR time horizon leaderboard: top 34 of 34 models (highest-effort row per model)
#ModelOrganisationMETR time horizon scoreCheapest input $/MSourceCaptured
1Claude Mythos Preview (Early)Anthropic1,045 min$10.00Epoch AI Benchmarking Hub2026-10-05
2GPT-5.4 (xhigh)OpenAI342 min$2.50Epoch AI Benchmarking Hub2026-10-05
3Gemini 3 Pro PreviewGoogle DeepMind224 min$2.00Epoch AI Benchmarking Hub2026-10-05
4GPT-5.1-Codex-MaxOpenAI224 min$1.25Epoch AI Benchmarking Hub2026-10-05
5GPT-5 (high)OpenAI203 min$1.25Epoch AI Benchmarking Hub2026-10-05
6claude-opus-4-1-20250805_16KAnthropic114 min—Epoch AI Benchmarking Hub2026-10-05
7grok-4-0709xAI110 min—Epoch AI Benchmarking Hub2026-10-05
8claude-opus-4-1-20250805Anthropic100 min—Epoch AI Benchmarking Hub2026-10-05
9Claude Opus 4Anthropic100 min$15.00Epoch AI Benchmarking Hub2026-10-05
10claude-sonnet-4-20250514_16KAnthropic75 min—Epoch AI Benchmarking Hub2026-10-05
11claude-3-7-sonnet-20250219Anthropic60 min—Epoch AI Benchmarking Hub2026-10-05
12Kimi K2 ThinkingMoonshot54 min$0.550Epoch AI Benchmarking Hub2026-10-05
13Gemini 2.5 Pro Preview (Jun 2025)Google DeepMind39 min$1.25Epoch AI Benchmarking Hub2026-10-05
14DeepSeek-R1 (May 2025)DeepSeek31 min$0.400Epoch AI Benchmarking Hub2026-10-05
15DeepSeek-R1DeepSeek27 min$0.400Epoch AI Benchmarking Hub2026-10-05
16DeepSeek-V3 (Mar 2025)DeepSeek23 min$0.200Epoch AI Benchmarking Hub2026-10-05
17Claude 3.5 Sonnet (Oct 2024)Anthropic21 min$3.00Epoch AI Benchmarking Hub2026-10-05
18o1-previewOpenAI20 min$15.00Epoch AI Benchmarking Hub2026-10-05
19DeepSeek-V3DeepSeek18 min$0.200Epoch AI Benchmarking Hub2026-10-05
20Claude 3.5 Sonnet (Jun 2024)Anthropic11 min$3.00Epoch AI Benchmarking Hub2026-10-05
21GPT-4o (Nov 2024)OpenAI9 min$2.50Epoch AI Benchmarking Hub2026-10-05
22GPT-4o (Aug 2024)OpenAI7 min$2.50Epoch AI Benchmarking Hub2026-10-05
23gpt-4-turbo-2024-04-09OpenAI7 min$10.00Epoch AI Benchmarking Hub2026-10-05
24GPT-4 Turbo Preview (January 2024)OpenAI5 min$10.00Epoch AI Benchmarking Hub2026-10-05
25GPT-4 (Mar 2023)OpenAI5 min$30.00Epoch AI Benchmarking Hub2026-10-05
26qwen2.5-72b-instructAlibaba5 min$0.120Epoch AI Benchmarking Hub2026-10-05
27GPT-4 Turbo Preview (Nov 2023)OpenAI4 min$10.00Epoch AI Benchmarking Hub2026-10-05
28GPT-4 (Jun 2023)OpenAI4 min$30.00Epoch AI Benchmarking Hub2026-10-05
29claude-3-opus-20240229Anthropic4 min$15.00Epoch AI Benchmarking Hub2026-10-05
30gpt-4-turboOpenAI4 min$10.00Epoch AI Benchmarking Hub2026-10-05
31qwen2-72b-instructAlibaba2 min$0.900Epoch AI Benchmarking Hub2026-10-05
32gpt-3.5-turbo-instructOpenAI1 min$1.50Epoch AI Benchmarking Hub2026-10-05
33davinci-002OpenAI0 min$2.00Epoch AI Benchmarking Hub2026-10-05
34gpt2-xlOpenAI0 min—Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to GPT-5.1-Codex-Max at $1.25 per million input tokens (score 224 min).

What METR time horizon measures

METR time horizon measures length of task (minutes) done at 50% success.

How to read these scores

Higher scores are better; the scale is specific to this benchmark, so compare models only on this board. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks