EBR-Bench leaderboard: which LLMs score highest

EBR-Bench (evidence-based reasoning) currently has scores for 19 models. The top score is 76.2% by GPT-6 Astra (max); the median model scores 30.5% and the lowest scores 2.4%.

EBR-Bench leaderboard: top 19 of 19 models (highest-effort row per model)
#ModelOrganisationEBR-Bench scoreCheapest input $/MSourceCaptured
1GPT-6 Astra (max)OpenAI76.2%$10.00Epoch AI Benchmarking Hub2026-10-05
2Claude Opus 5.5 (max)Anthropic71.4%$4.00Epoch AI Benchmarking Hub2026-10-05
3Claude Fable 5.1 (max)Anthropic57.1%$10.00Epoch AI Benchmarking Hub2026-10-05
4GPT-6.1 Sol (max)OpenAI54.3%$2.00Epoch AI Benchmarking Hub2026-10-05
5GPT-6 Sol (max)OpenAI53.3%$2.00Epoch AI Benchmarking Hub2026-10-05
6Claude Opus 5 (max)Anthropic45.7%$5.00Epoch AI Benchmarking Hub2026-10-05
7GPT-5.6 Sol (max)OpenAI44.8%$2.00Epoch AI Benchmarking Hub2026-10-05
8Claude Fable 5 (max)Anthropic39.5%$10.00Epoch AI Benchmarking Hub2026-10-05
9GPT-5.5 (xhigh)OpenAI34.3%$5.00Epoch AI Benchmarking Hub2026-10-05
10Grok 4.6 (xhigh)xAI30.5%$1.25Epoch AI Benchmarking Hub2026-10-05
11GPT-5.4 (xhigh)OpenAI25.4%$2.50Epoch AI Benchmarking Hub2026-10-05
12GPT-5.2 (xhigh)OpenAI23.0%$1.75Epoch AI Benchmarking Hub2026-10-05
13Claude Opus 4.7 (max)Anthropic19.0%$5.00Epoch AI Benchmarking Hub2026-10-05
14Claude Opus 4.5 (128k thinking)Anthropic14.3%$5.00Epoch AI Benchmarking Hub2026-10-05
15Claude Opus 4.6 (max)Anthropic12.7%$5.00Epoch AI Benchmarking Hub2026-10-05
16GPT-5 (high)OpenAI12.7%$1.25Epoch AI Benchmarking Hub2026-10-05
17Qwen3.7 MaxAlibaba9.5%$1.25Epoch AI Benchmarking Hub2026-10-05
18claude-opus-4-1-20250805Anthropic7.9%—Epoch AI Benchmarking Hub2026-10-05
19Kimi K2.6Moonshot2.4%$0.650Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to Grok 4.6 (xhigh) at $1.25 per million input tokens (score 30.5%).

What EBR-Bench measures

EBR-Bench measures evidence-based reasoning.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks