LMCA leaderboard: which LLMs score highest

LMCA (language-model capability assessment) currently has scores for 111 models. The top score is 68.2% by Claude Opus 5.5 (max); the median model scores 29.5% and the lowest scores 2.8%.

LMCA leaderboard: top 50 of 111 models (highest-effort row per model)
#ModelOrganisationLMCA scoreCheapest input $/MSourceCaptured
1Claude Opus 5.5 (max)Anthropic68.2%$4.00Epoch AI Benchmarking Hub2026-10-05
2Claude Opus 5 (max)Anthropic63.3%$5.00Epoch AI Benchmarking Hub2026-10-05
3Claude Fable 5 (max)Anthropic60.3%$10.00Epoch AI Benchmarking Hub2026-10-05
4GPT-5.6 Sol (pro, max)OpenAI59.2%$2.00Epoch AI Benchmarking Hub2026-10-05
5GPT-6 Sol (max)OpenAI59.1%$2.00Epoch AI Benchmarking Hub2026-10-05
6GPT-5.6 Sol (max)OpenAI58.4%$2.00Epoch AI Benchmarking Hub2026-10-05
7Claude Opus 4.6 (max)Anthropic55.8%$5.00Epoch AI Benchmarking Hub2026-10-05
8GPT-5.6 Terra (max)OpenAI55.0%$2.00Epoch AI Benchmarking Hub2026-10-05
9GPT-5.5 (xhigh)OpenAI54.3%$5.00Epoch AI Benchmarking Hub2026-10-05
10GPT-5.5 Pro (xhigh)OpenAI53.9%$30.00Epoch AI Benchmarking Hub2026-10-05
11Kimi K3 (max)Moonshot52.7%$1.29Epoch AI Benchmarking Hub2026-10-05
12Claude Opus 4.7 (max)Anthropic52.2%$5.00Epoch AI Benchmarking Hub2026-10-05
13GPT-5.4 (xhigh)OpenAI52.0%$2.50Epoch AI Benchmarking Hub2026-10-05
14Gemini 3.7 Flash (high)Google DeepMind50.4%$0.750Epoch AI Benchmarking Hub2026-10-05
15Muse Spark 1.1 (high)Meta AI49.9%$1.25Epoch AI Benchmarking Hub2026-10-05
16Claude Sonnet 5 (max)Anthropic49.3%$2.00Epoch AI Benchmarking Hub2026-10-05
17GPT-5.6 Luna (max)OpenAI48.5%$0.200Epoch AI Benchmarking Hub2026-10-05
18Grok 4.6 (xhigh)xAI48.5%$1.25Epoch AI Benchmarking Hub2026-10-05
19Muse Spark 1.2 (xhigh)Meta AI48.4%$1.25Epoch AI Benchmarking Hub2026-10-05
20Claude Sonnet 4.6 (max)Anthropic46.5%$3.00Epoch AI Benchmarking Hub2026-10-05
21Qwen3.8 Max (xhigh)Alibaba46.2%$1.65Epoch AI Benchmarking Hub2026-10-05
22Grok 4.5 (high)xAI45.2%$2.00Epoch AI Benchmarking Hub2026-10-05
23GPT-6 Luna (max)OpenAI44.5%$0.100Epoch AI Benchmarking Hub2026-10-05
24Qwen3.7 MaxAlibaba44.0%$1.25Epoch AI Benchmarking Hub2026-10-05
25GPT-5.2 (xhigh)OpenAI43.9%$1.75Epoch AI Benchmarking Hub2026-10-05
26GPT-5.1 (high)OpenAI43.9%$1.25Epoch AI Benchmarking Hub2026-10-05
27Qwen 3.6 Max (Preview)Alibaba42.5%—Epoch AI Benchmarking Hub2026-10-05
28DeepSeek V4 Flash 0731 (max)DeepSeek41.7%$0.015Epoch AI Benchmarking Hub2026-10-05
29DeepSeek v4 Pro (max)DeepSeek41.2%$0.435Epoch AI Benchmarking Hub2026-10-05
30GPT-5.4 mini (xhigh)OpenAI40.8%$0.750Epoch AI Benchmarking Hub2026-10-05
31GPT-5 (high)OpenAI40.0%$1.25Epoch AI Benchmarking Hub2026-10-05
32o3 (high)OpenAI39.7%$2.00Epoch AI Benchmarking Hub2026-10-05
33Gemma 4 31B ITGoogle DeepMind39.2%$0.100Epoch AI Benchmarking Hub2026-10-05
34grok-4-20xAI38.7%$1.25Epoch AI Benchmarking Hub2026-10-05
35o3-pro-2025-06-10 (high)OpenAI38.5%$20.00Epoch AI Benchmarking Hub2026-10-05
36grok-4.3 (high)xAI38.3%$1.25Epoch AI Benchmarking Hub2026-10-05
37Qwen3.5 397B-A17BAlibaba37.9%$0.164Epoch AI Benchmarking Hub2026-10-05
38Inkling (xhigh)Thinking Machines37.6%$0.950Epoch AI Benchmarking Hub2026-10-05
39Qwen3.7 PlusAlibaba37.6%$0.282Epoch AI Benchmarking Hub2026-10-05
40Claude Opus 4Anthropic37.4%$15.00Epoch AI Benchmarking Hub2026-10-05
41Kimi K2.6Moonshot37.3%$0.650Epoch AI Benchmarking Hub2026-10-05
42Claude Opus 4.1 (unknown thinking)Anthropic37.1%$15.00Epoch AI Benchmarking Hub2026-10-05
43nemotron-3-ultraNvidia36.9%$0.500Epoch AI Benchmarking Hub2026-10-05
44GPT-5.4 nano (xhigh)OpenAI36.9%$0.200Epoch AI Benchmarking Hub2026-10-05
45Qwen3.5 PlusAlibaba36.4%—Epoch AI Benchmarking Hub2026-10-05
46DeepSeek v4 Flash (max)DeepSeek35.9%$0.090Epoch AI Benchmarking Hub2026-10-05
47Qwen3.6 27BAlibaba34.5%$0.150Epoch AI Benchmarking Hub2026-10-05
48GPT-5 mini (high)OpenAI34.2%$0.250Epoch AI Benchmarking Hub2026-10-05
49Qwen3.5 27BAlibaba34.0%$0.195Epoch AI Benchmarking Hub2026-10-05
50MiniMax-M3MiniMax33.7%$0.230Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to GPT-5.6 Sol (pro, max) at $2.00 per million input tokens (score 59.2%).

What LMCA measures

LMCA measures language-model capability assessment.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks