CritPt leaderboard: which LLMs score highest

CritPt (unpublished research-level physics problems) currently has scores for 111 models. The top score is 32.3% by GPT-5.6 Sol (max); the median model scores 2.9% and the lowest scores 0.0%.

CritPt leaderboard: top 50 of 111 models (highest-effort row per model)
#ModelOrganisationCritPt scoreCheapest input $/MSourceCaptured
1GPT-5.6 Sol (max)OpenAI32.3%$2.00Epoch AI Benchmarking Hub2026-10-05
2Claude Opus 5.5 (max)Anthropic31.7%$4.00Epoch AI Benchmarking Hub2026-10-05
3GPT-6 Astra (max)OpenAI31.7%$10.00Epoch AI Benchmarking Hub2026-10-05
4GPT-6.1 Sol (max)OpenAI31.7%$2.00Epoch AI Benchmarking Hub2026-10-05
5Claude Sonnet 5.5 (max)Anthropic31.4%$2.00Epoch AI Benchmarking Hub2026-10-05
6GPT-6 Sol (max)OpenAI30.9%$2.00Epoch AI Benchmarking Hub2026-10-05
7GPT-5.5 Pro (xhigh)OpenAI30.6%$30.00Epoch AI Benchmarking Hub2026-10-05
8GPT-5.4 Pro (xhigh)OpenAI30.0%$30.00Epoch AI Benchmarking Hub2026-10-05
9GPT-5.6 Terra (max)OpenAI30.0%$2.00Epoch AI Benchmarking Hub2026-10-05
10Claude Fable 5.1 (max)Anthropic29.7%$10.00Epoch AI Benchmarking Hub2026-10-05
11Claude Opus 5 (max)Anthropic29.1%$5.00Epoch AI Benchmarking Hub2026-10-05
12Claude Fable 5 (max)Anthropic28.6%$10.00Epoch AI Benchmarking Hub2026-10-05
13GPT-5.5 (xhigh)OpenAI27.1%$5.00Epoch AI Benchmarking Hub2026-10-05
14Gemini 3 Deep ThinkGoogle DeepMind25.7%—Epoch AI Benchmarking Hub2026-10-05
15Muse Spark 1.3 (max)Meta AI24.9%$1.25Epoch AI Benchmarking Hub2026-10-05
16GPT-5.4 (xhigh)OpenAI23.4%$2.50Epoch AI Benchmarking Hub2026-10-05
17Kimi K3 (max)Moonshot23.4%$1.29Epoch AI Benchmarking Hub2026-10-05
18GPT-5.6 Luna (max)OpenAI20.6%$0.200Epoch AI Benchmarking Hub2026-10-05
19Grok 4.6 (xhigh)xAI19.7%$1.25Epoch AI Benchmarking Hub2026-10-05
20GPT-6 Luna (max)OpenAI19.4%$0.100Epoch AI Benchmarking Hub2026-10-05
21GLM-5.3 (max)Z.ai (Zhipu AI)19.1%$0.070Epoch AI Benchmarking Hub2026-10-05
22Gemini 3.8 Flash (high)Google DeepMind18.3%$0.750Epoch AI Benchmarking Hub2026-10-05
23DeepSeek V4 Pro 0813 (max)DeepSeek18.0%$0.660Epoch AI Benchmarking Hub2026-10-05
24Grok 4.7 (xhigh)xAI17.7%$2.00Epoch AI Benchmarking Hub2026-10-05
25Muse Spark 1.2 (xhigh)Meta AI17.7%$1.25Epoch AI Benchmarking Hub2026-10-05
26Claude Sonnet 5 (max)Anthropic16.9%$2.00Epoch AI Benchmarking Hub2026-10-05
27DeepSeek V4 Flash 0731 (max)DeepSeek16.6%$0.015Epoch AI Benchmarking Hub2026-10-05
28Grok 4.5 (high)xAI15.4%$2.00Epoch AI Benchmarking Hub2026-10-05
29deepseek-v4.1-flash-maxDeepSeek14.3%—Epoch AI Benchmarking Hub2026-10-05
30Gemini 3.7 Flash (high)Google DeepMind14.3%$0.750Epoch AI Benchmarking Hub2026-10-05
31Qwen3.7 MaxAlibaba13.4%$1.25Epoch AI Benchmarking Hub2026-10-05
32DeepSeek v4 Pro (max)DeepSeek12.9%$0.435Epoch AI Benchmarking Hub2026-10-05
33GPT-5 (high)OpenAI12.6%$1.25Epoch AI Benchmarking Hub2026-10-05
34Claude Opus 4.7 (max)Anthropic12.0%$5.00Epoch AI Benchmarking Hub2026-10-05
35Muse SparkMeta AI11.3%—Epoch AI Benchmarking Hub2026-10-05
36GPT-5.4 mini (xhigh)OpenAI10.0%$0.750Epoch AI Benchmarking Hub2026-10-05
37Kimi K2.7 CodeMoonshot10.0%$0.671Epoch AI Benchmarking Hub2026-10-05
38GPT-5.4 nano (xhigh)OpenAI9.3%$0.200Epoch AI Benchmarking Hub2026-10-05
39grok-build-0.1xAI9.1%$1.00Epoch AI Benchmarking Hub2026-10-05
40Qwen3.7 PlusAlibaba9.1%$0.282Epoch AI Benchmarking Hub2026-10-05
41grok-4.3 (high)xAI8.0%$1.25Epoch AI Benchmarking Hub2026-10-05
42Kimi K2.6Moonshot8.0%$0.650Epoch AI Benchmarking Hub2026-10-05
43DeepSeek v4 Flash (max)DeepSeek7.1%$0.090Epoch AI Benchmarking Hub2026-10-05
44Gemini 3 Pro PreviewGoogle DeepMind6.9%$2.00Epoch AI Benchmarking Hub2026-10-05
45Inkling (xhigh)Thinking Machines5.4%$0.950Epoch AI Benchmarking Hub2026-10-05
46Qwen3.8-27B (XHigh)Alibaba5.4%$0.400Epoch AI Benchmarking Hub2026-10-05
47GLM-5.1Z.ai (Zhipu AI)4.6%$1.05Epoch AI Benchmarking Hub2026-10-05
48mimo-v2.5-proXiaomi Corp4.0%$0.435Epoch AI Benchmarking Hub2026-10-05
49mimo-v2.5Xiaomi Corp3.7%$0.140Epoch AI Benchmarking Hub2026-10-05
50MiniMax-M3MiniMax3.7%$0.230Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to GPT-5.6 Sol (max) at $2.00 per million input tokens (score 32.3%).

What CritPt measures

CritPt measures unpublished research-level physics problems.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks