LiveBench Reasoning leaderboard: which LLMs score highest

LiveBench Reasoning (LiveBench release 2026_06_25) currently has scores for 17 models. The top score is 89.7% by claude-opus-4-8-xhigh-effort; the median model scores 81.1% and the lowest scores 70.3%.

LiveBench Reasoning leaderboard: top 17 of 17 models (highest-effort row per model)
#ModelOrganisationLiveBench Reasoning scoreCheapest input $/MSourceCaptured
1claude-opus-4-8-xhigh-effortAnthropic89.7%$5.00LiveBench 2026_06_25 official CSV2026-10-05
2GPT-5.5 (xhigh)OpenAI89.7%$5.00LiveBench 2026_06_25 official CSV2026-10-05
3GPT-5.4 (xhigh)OpenAI88.1%$2.50LiveBench 2026_06_25 official CSV2026-10-05
4gemini-3.1-pro-preview-highGoogle84.0%$2.00LiveBench 2026_06_25 official CSV2026-10-05
5Qwen3.7 MaxAlibaba83.3%$1.25LiveBench 2026_06_25 official CSV2026-10-05
6gpt-5.2-2025-12-11-highOpenAI83.2%$1.75LiveBench 2026_06_25 official CSV2026-10-05
7Kimi K2.7 CodeMoonshot82.8%$0.671LiveBench 2026_06_25 official CSV2026-10-05
8gemini-3.5-flash-highGoogle82.0%$1.50LiveBench 2026_06_25 official CSV2026-10-05
9GPT-5.4 nano (xhigh)OpenAI81.1%$0.200LiveBench 2026_06_25 official CSV2026-10-05
10claude-opus-4-5-20251101-thinking-64k-high-effortAnthropic80.1%$5.00LiveBench 2026_06_25 official CSV2026-10-05
11kimi-k2.6-thinkingMoonshot AI79.4%$0.650LiveBench 2026_06_25 official CSV2026-10-05
12GPT-5.2 CodexOpenAI77.7%$1.75LiveBench 2026_06_25 official CSV2026-10-05
13grok-build-0.1xAI76.4%$1.00LiveBench 2026_06_25 official CSV2026-10-05
14qwen3.6-plusAlibaba75.8%$0.325LiveBench 2026_06_25 official CSV2026-10-05
15MiniMax-M3MiniMax74.5%$0.230LiveBench 2026_06_25 official CSV2026-10-05
16GPT-5.4 mini (xhigh)OpenAI71.3%$0.750LiveBench 2026_06_25 official CSV2026-10-05
17Qwen3.6 27BAlibaba70.3%$0.150LiveBench 2026_06_25 official CSV2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to GPT-5.4 nano (xhigh) at $0.200 per million input tokens (score 81.1%).

What LiveBench Reasoning measures

LiveBench refreshes its questions regularly from recent material to limit training-data contamination, and scores answers automatically against ground truth. This is its reasoning category.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

LiveBench 2026_06_25 official CSV. Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks