DTBench leaderboard: which LLMs score highest

DTBench (dynamic tool-use benchmark) currently has scores for 149 models. The top score is 98.9% by Claude Opus 5.5 (max); the median model scores 69.3% and the lowest scores 41.6%.

DTBench leaderboard: top 50 of 149 models (highest-effort row per model)
#ModelOrganisationDTBench scoreCheapest input $/MSourceCaptured
1Claude Opus 5.5 (max)Anthropic98.9%$4.00Epoch AI Benchmarking Hub2026-10-05
2Claude Fable 5 (max)Anthropic98.4%$10.00Epoch AI Benchmarking Hub2026-10-05
3Claude Opus 5 (max)Anthropic97.6%$5.00Epoch AI Benchmarking Hub2026-10-05
4GPT-6 Sol (max)OpenAI97.3%$2.00Epoch AI Benchmarking Hub2026-10-05
5Grok 4.6 (xhigh)xAI97.3%$1.25Epoch AI Benchmarking Hub2026-10-05
6Gemini 3.7 Flash (high)Google DeepMind96.8%$0.750Epoch AI Benchmarking Hub2026-10-05
7Grok 4.5 (high)xAI96.5%$2.00Epoch AI Benchmarking Hub2026-10-05
8GPT-5.5 (xhigh)OpenAI96.0%$5.00Epoch AI Benchmarking Hub2026-10-05
9GPT-5.5 Pro (xhigh)OpenAI96.0%$30.00Epoch AI Benchmarking Hub2026-10-05
10GPT-5.6 Sol (pro, max)OpenAI96.0%$2.00Epoch AI Benchmarking Hub2026-10-05
11GPT-5.6 Sol (max)OpenAI95.5%$2.00Epoch AI Benchmarking Hub2026-10-05
12Muse Spark 1.2 (xhigh)Meta AI94.7%$1.25Epoch AI Benchmarking Hub2026-10-05
13Claude Opus 4.7 (max)Anthropic94.7%$5.00Epoch AI Benchmarking Hub2026-10-05
14GPT-5.4 (xhigh)OpenAI94.4%$2.50Epoch AI Benchmarking Hub2026-10-05
15Muse Spark 1.1 (high)Meta AI94.4%$1.25Epoch AI Benchmarking Hub2026-10-05
16GPT-5.6 Terra (max)OpenAI93.3%$2.00Epoch AI Benchmarking Hub2026-10-05
17Claude Sonnet 5 (max)Anthropic92.5%$2.00Epoch AI Benchmarking Hub2026-10-05
18Qwen3.7 MaxAlibaba92.3%$1.25Epoch AI Benchmarking Hub2026-10-05
19Qwen3.8 Max (xhigh)Alibaba92.0%$1.65Epoch AI Benchmarking Hub2026-10-05
20Claude Opus 4.6 (max)Anthropic91.2%$5.00Epoch AI Benchmarking Hub2026-10-05
21Kimi K3 (max)Moonshot91.2%$1.29Epoch AI Benchmarking Hub2026-10-05
22DeepSeek V4 Flash 0731 (max)DeepSeek90.9%$0.015Epoch AI Benchmarking Hub2026-10-05
23GPT-5.2 (xhigh)OpenAI90.9%$1.75Epoch AI Benchmarking Hub2026-10-05
24Kimi K2.6Moonshot90.9%$0.650Epoch AI Benchmarking Hub2026-10-05
25DeepSeek v4 Pro (max)DeepSeek90.7%$0.435Epoch AI Benchmarking Hub2026-10-05
26GPT-5 (high)OpenAI90.7%$1.25Epoch AI Benchmarking Hub2026-10-05
27grok-4.3 (high)xAI90.7%$1.25Epoch AI Benchmarking Hub2026-10-05
28GPT-5.1 (high)OpenAI90.1%$1.25Epoch AI Benchmarking Hub2026-10-05
29GPT-6 Luna (max)OpenAI90.1%$0.100Epoch AI Benchmarking Hub2026-10-05
30grok-4-20xAI90.1%$1.25Epoch AI Benchmarking Hub2026-10-05
31nemotron-3-ultraNvidia90.1%$0.500Epoch AI Benchmarking Hub2026-10-05
32Claude Sonnet 4.6 (max)Anthropic89.9%$3.00Epoch AI Benchmarking Hub2026-10-05
33GPT-5.6 Luna (max)OpenAI88.8%$0.200Epoch AI Benchmarking Hub2026-10-05
34grok-4-1-fast-reasoningxAI87.7%$0.200Epoch AI Benchmarking Hub2026-10-05
35Inkling (xhigh)Thinking Machines87.5%$0.950Epoch AI Benchmarking Hub2026-10-05
36Qwen3.5 397B-A17BAlibaba87.5%$0.164Epoch AI Benchmarking Hub2026-10-05
37Qwen 3.6 Max (Preview)Alibaba87.2%—Epoch AI Benchmarking Hub2026-10-05
38o3-pro-2025-06-10 (high)OpenAI86.9%$20.00Epoch AI Benchmarking Hub2026-10-05
39DeepSeek v4 Flash (max)DeepSeek86.4%$0.090Epoch AI Benchmarking Hub2026-10-05
40DeepSeek-V3.2 (Thinking; Novita)DeepSeek85.6%$0.259Epoch AI Benchmarking Hub2026-10-05
41o3 (high)OpenAI84.8%$2.00Epoch AI Benchmarking Hub2026-10-05
42mimo-v2.5-proXiaomi Corp84.5%$0.435Epoch AI Benchmarking Hub2026-10-05
43Qwen3.7 PlusAlibaba84.0%$0.282Epoch AI Benchmarking Hub2026-10-05
44Qwen3.5 FlashAlibaba82.9%$0.065Epoch AI Benchmarking Hub2026-10-05
45Gemma 4 31B ITGoogle DeepMind82.7%$0.100Epoch AI Benchmarking Hub2026-10-05
46grok-4-fastxAI82.7%—Epoch AI Benchmarking Hub2026-10-05
47Qwen3.5 27BAlibaba82.4%$0.195Epoch AI Benchmarking Hub2026-10-05
48Qwen3-Max-InstructAlibaba82.1%—Epoch AI Benchmarking Hub2026-10-05
49Qwen 3.6 Plus (2026-04-02)Alibaba81.9%—Epoch AI Benchmarking Hub2026-10-05
50Claude Opus 4Anthropic81.6%$15.00Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to Gemini 3.7 Flash (high) at $0.750 per million input tokens (score 96.8%).

What DTBench measures

DTBench measures dynamic tool-use benchmark.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks