GPQA Diamond leaderboard: which LLMs score highest

GPQA Diamond (expert-level science questions) currently has scores for 274 models. The top score is 95.8% by GPT-6 Astra (max); the median model scores 71.6% and the lowest scores 9.3%.

GPQA Diamond leaderboard: top 50 of 274 models (highest-effort row per model)
#ModelOrganisationGPQA Diamond scoreCheapest input $/MSourceCaptured
1GPT-6 Astra (max)OpenAI95.8%$10.00Epoch AI Benchmarking Hub2026-10-05
2Claude Sonnet 5.5 (max)Anthropic95.6%$2.00Epoch AI Benchmarking Hub2026-10-05
3Gemini 3.8 Flash (high)Google DeepMind95.4%$0.750Epoch AI Benchmarking Hub2026-10-05
4GPT-6.1 Sol (max)OpenAI95.4%$2.00Epoch AI Benchmarking Hub2026-10-05
5Gemini 3.7 Flash (high)Google DeepMind94.8%$0.750Epoch AI Benchmarking Hub2026-10-05
6GPT-5.4 Pro (xhigh)OpenAI94.6%$30.00Epoch AI Benchmarking Hub2026-10-05
7GPT-6 Sol (max)OpenAI94.3%$2.00Epoch AI Benchmarking Hub2026-10-05
8Claude Opus 5 (max)Anthropic93.9%$5.00Epoch AI Benchmarking Hub2026-10-05
9GPT-5.6 Sol (max)OpenAI93.5%$2.00Epoch AI Benchmarking Hub2026-10-05
10Grok 4.5 (high)xAI93.4%$2.00Epoch AI Benchmarking Hub2026-10-05
11Darwin-397B-ZTCFINAL-Bench93.4%—HuggingFace GPQA Diamond leaderboard2026-10-05
12GPT-5.6 Terra (max)OpenAI93.3%$2.00Epoch AI Benchmarking Hub2026-10-05
13GPT-5.4 (xhigh)OpenAI93.3%$2.50Epoch AI Benchmarking Hub2026-10-05
14Grok 4.6 (xhigh)xAI93.2%$1.25Epoch AI Benchmarking Hub2026-10-05
15Kimi K3 (max)Moonshot93.1%$1.29Epoch AI Benchmarking Hub2026-10-05
16Ornith-1.5-397Bornith-ai92.8%—HuggingFace GPQA Diamond leaderboard2026-10-05
17Grok 4.7 (xhigh)xAI92.7%$2.00Epoch AI Benchmarking Hub2026-10-05
18Qwen3.8 Max (xhigh)Alibaba92.7%$1.65Epoch AI Benchmarking Hub2026-10-05
19Gemini 3 Pro PreviewGoogle DeepMind92.6%$2.00Epoch AI Benchmarking Hub2026-10-05
20Qwen3.8-2.4T-A95BQwen92.6%$2.00HuggingFace GPQA Diamond leaderboard2026-10-05
21Hy4-previewtencent92.3%$0.751HuggingFace GPQA Diamond leaderboard2026-10-05
22Qwen3.8 Max (0902) (xhigh)Alibaba92.3%$1.65Epoch AI Benchmarking Hub2026-10-05
23Qwen3.8-Flash-NextQwen91.7%—HuggingFace GPQA Diamond leaderboard2026-10-05
24DeepSeek V4 Pro 0813 (max)DeepSeek91.7%$0.660Epoch AI Benchmarking Hub2026-10-05
25GPT-5.6 Luna (max)OpenAI91.6%$0.200Epoch AI Benchmarking Hub2026-10-05
26GPT-5.2 (xhigh)OpenAI91.4%$1.75Epoch AI Benchmarking Hub2026-10-05
27DeepSeek V4 Flash 0731 (max)DeepSeek91.0%$0.015Epoch AI Benchmarking Hub2026-10-05
28GLM-5.3 (max)Z.ai (Zhipu AI)90.9%$0.070Epoch AI Benchmarking Hub2026-10-05
29MiniMax-M3MiniMax90.9%$0.230Epoch AI Benchmarking Hub2026-10-05
30Qwen3.7 MaxAlibaba90.9%$1.25Epoch AI Benchmarking Hub2026-10-05
31Darwin-398B-JGOSFINAL-Bench90.9%—HuggingFace GPQA Diamond leaderboard2026-10-05
32Claude Opus 5.5 (max)Anthropic90.6%$4.00Epoch AI Benchmarking Hub2026-10-05
33Kimi K2.6Moonshot90.5%$0.650HuggingFace GPQA Diamond leaderboard2026-10-05
34GPT-6 Luna (max)OpenAI90.5%$0.100Epoch AI Benchmarking Hub2026-10-05
35Hy3tencent90.4%$0.083HuggingFace GPQA Diamond leaderboard2026-10-05
36GLM-5.3-Flash (max)Z.ai (Zhipu AI)90.2%$0.110Epoch AI Benchmarking Hub2026-10-05
37Muse SparkMeta AI89.8%—Epoch AI Benchmarking Hub2026-10-05
38DeepSeek v4 Pro (max)DeepSeek89.6%$0.435Epoch AI Benchmarking Hub2026-10-05
39grok-4.20-0309-reasoningxAI89.3%$1.25Epoch AI Benchmarking Hub2026-10-05
40Ornith-1.5-35B-A3Bornith-ai89.2%—HuggingFace GPQA Diamond leaderboard2026-10-05
41grok-4.3 (high)xAI88.8%$1.25Epoch AI Benchmarking Hub2026-10-05
42Inkling Small (xhigh)Thinking Machines88.5%$0.450Epoch AI Benchmarking Hub2026-10-05
43Darwin-36B-OpusFINAL-Bench88.4%—HuggingFace GPQA Diamond leaderboard2026-10-05
44Qwen3.5 397B-A17BAlibaba88.4%$0.164HuggingFace GPQA Diamond leaderboard2026-10-05
45Claude Opus 4.6 (max)Anthropic88.4%$5.00Epoch AI Benchmarking Hub2026-10-05
46Qwen 3.6 Plus (2026-04-02)Alibaba88.4%—Epoch AI Benchmarking Hub2026-10-05
47Ring-2.6-1TinclusionAI88.3%—HuggingFace GPQA Diamond leaderboard2026-10-05
48Inkling (xhigh)Thinking Machines88.3%$0.950Epoch AI Benchmarking Hub2026-10-05
49NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4nvidia87.9%—HuggingFace GPQA Diamond leaderboard2026-10-05
50Kimi K2.7 CodeMoonshot87.9%$0.671Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to Gemini 3.8 Flash (high) at $0.750 per million input tokens (score 95.4%).

What GPQA Diamond measures

GPQA Diamond is the hardest subset of GPQA: 198 multiple-choice questions in biology, physics and chemistry written and checked by domain experts. The questions are designed to be "Google-proof": skilled non-experts with web access still do badly, so a high score points to real scientific reasoning rather than lookup.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0) and HuggingFace GPQA Diamond leaderboard. Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks