GDP.pdf leaderboard: which LLMs score highest

GDP.pdf (economically valuable document tasks) currently has scores for 22 models. The top score is 30.7% by GPT-5.6 Sol (max); the median model scores 22.0% and the lowest scores 2.0%.

GDP.pdf leaderboard: top 22 of 22 models (highest-effort row per model)
#ModelOrganisationGDP.pdf scoreCheapest input $/MSourceCaptured
1GPT-5.6 Sol (max)OpenAI30.7%$2.00Epoch AI Benchmarking Hub2026-10-05
2Claude Opus 5.5 (max)Anthropic30.6%$4.00Epoch AI Benchmarking Hub2026-10-05
3Claude Fable 5 (max)Anthropic29.8%$10.00Epoch AI Benchmarking Hub2026-10-05
4Claude Fable 5.1 (max)Anthropic27.6%$10.00Epoch AI Benchmarking Hub2026-10-05
5GPT-6 Sol (max)OpenAI26.4%$2.00Epoch AI Benchmarking Hub2026-10-05
6GPT-5.5 (xhigh)OpenAI26.0%$5.00Epoch AI Benchmarking Hub2026-10-05
7Claude Opus 5 (max)Anthropic24.0%$5.00Epoch AI Benchmarking Hub2026-10-05
8Gemini 3.7 Flash (high)Google DeepMind23.8%$0.750Epoch AI Benchmarking Hub2026-10-05
9Gemini 3.8 Flash (high)Google DeepMind23.2%$0.750Epoch AI Benchmarking Hub2026-10-05
10Qwen3.8 Max (xhigh)Alibaba23.2%$1.65Epoch AI Benchmarking Hub2026-10-05
11GPT-6 Luna (max)OpenAI23.0%$0.100Epoch AI Benchmarking Hub2026-10-05
12Claude Opus 4.7 (max)Anthropic21.0%$5.00Epoch AI Benchmarking Hub2026-10-05
13Kimi K3 (max)Moonshot19.0%$1.29Epoch AI Benchmarking Hub2026-10-05
14Claude Sonnet 4.6 (max)Anthropic18.0%$3.00Epoch AI Benchmarking Hub2026-10-05
15Grok 4.6 (xhigh)xAI17.2%$1.25Epoch AI Benchmarking Hub2026-10-05
16Qwen3.8 Flash (xhigh)Alibaba16.6%$0.090Epoch AI Benchmarking Hub2026-10-05
17Muse Spark 1.2 (xhigh)Meta AI16.0%$1.25Epoch AI Benchmarking Hub2026-10-05
18GLM-5.3-Flash (max)Z.ai (Zhipu AI)14.0%$0.110Epoch AI Benchmarking Hub2026-10-05
19Grok 4.5 (high)xAI14.0%$2.00Epoch AI Benchmarking Hub2026-10-05
20Kimi K2.6Moonshot12.0%$0.650Epoch AI Benchmarking Hub2026-10-05
21grok-4.3 (high)xAI8.0%$1.25Epoch AI Benchmarking Hub2026-10-05
22Nova 2.0 Pro Preview (unknown)Amazon2.0%—Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to Gemini 3.7 Flash (high) at $0.750 per million input tokens (score 23.8%).

What GDP.pdf measures

GDP.pdf measures economically valuable document tasks.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks