LiveBench Language leaderboard: which LLMs score highest

LiveBench Language (LiveBench release 2026_06_25) currently has scores for 17 models. The top score is 87.4% by GPT-5.5 (xhigh); the median model scores 77.9% and the lowest scores 62.5%.

LiveBench Language leaderboard: top 17 of 17 models (highest-effort row per model)
#ModelOrganisationLiveBench Language scoreCheapest input $/MSourceCaptured
1GPT-5.5 (xhigh)OpenAI87.4%$5.00LiveBench 2026_06_25 official CSV2026-10-05
2gemini-3.1-pro-preview-highGoogle85.4%$2.00LiveBench 2026_06_25 official CSV2026-10-05
3gemini-3.5-flash-highGoogle84.6%$1.50LiveBench 2026_06_25 official CSV2026-10-05
4GPT-5.4 (xhigh)OpenAI82.6%$2.50LiveBench 2026_06_25 official CSV2026-10-05
5claude-opus-4-8-xhigh-effortAnthropic81.4%$5.00LiveBench 2026_06_25 official CSV2026-10-05
6claude-opus-4-5-20251101-thinking-64k-high-effortAnthropic81.3%$5.00LiveBench 2026_06_25 official CSV2026-10-05
7gpt-5.2-2025-12-11-highOpenAI79.8%$1.75LiveBench 2026_06_25 official CSV2026-10-05
8Qwen3.7 MaxAlibaba79.7%$1.25LiveBench 2026_06_25 official CSV2026-10-05
9Kimi K2.7 CodeMoonshot77.9%$0.671LiveBench 2026_06_25 official CSV2026-10-05
10MiniMax-M3MiniMax76.8%$0.230LiveBench 2026_06_25 official CSV2026-10-05
11kimi-k2.6-thinkingMoonshot AI75.1%$0.650LiveBench 2026_06_25 official CSV2026-10-05
12qwen3.6-plusAlibaba75.0%$0.325LiveBench 2026_06_25 official CSV2026-10-05
13GPT-5.2 CodexOpenAI73.7%$1.75LiveBench 2026_06_25 official CSV2026-10-05
14grok-build-0.1xAI72.5%$1.00LiveBench 2026_06_25 official CSV2026-10-05
15GPT-5.4 mini (xhigh)OpenAI71.0%$0.750LiveBench 2026_06_25 official CSV2026-10-05
16Qwen3.6 27BAlibaba63.3%$0.150LiveBench 2026_06_25 official CSV2026-10-05
17GPT-5.4 nano (xhigh)OpenAI62.5%$0.200LiveBench 2026_06_25 official CSV2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to MiniMax-M3 at $0.230 per million input tokens (score 76.8%).

What LiveBench Language measures

LiveBench Language measures LiveBench release 2026_06_25.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

LiveBench 2026_06_25 official CSV. Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks