GPQA Diamond leaderboard: which LLMs score highest
GPQA Diamond (expert-level science questions) currently has scores for 274 models. The top score is 95.8% by GPT-6 Astra (max); the median model scores 71.6% and the lowest scores 9.3%.
| # | Model | Organisation | GPQA Diamond score | Cheapest input $/M | Source | Captured |
|---|---|---|---|---|---|---|
| 1 | GPT-6 Astra (max) | OpenAI | 95.8% | $10.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 2 | Claude Sonnet 5.5 (max) | Anthropic | 95.6% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 3 | Gemini 3.8 Flash (high) | Google DeepMind | 95.4% | $0.750 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 4 | GPT-6.1 Sol (max) | OpenAI | 95.4% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 5 | Gemini 3.7 Flash (high) | Google DeepMind | 94.8% | $0.750 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 6 | GPT-5.4 Pro (xhigh) | OpenAI | 94.6% | $30.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 7 | GPT-6 Sol (max) | OpenAI | 94.3% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 8 | Claude Opus 5 (max) | Anthropic | 93.9% | $5.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 9 | GPT-5.6 Sol (max) | OpenAI | 93.5% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 10 | Grok 4.5 (high) | xAI | 93.4% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 11 | Darwin-397B-ZTC | FINAL-Bench | 93.4% | — | HuggingFace GPQA Diamond leaderboard | 2026-10-05 |
| 12 | GPT-5.6 Terra (max) | OpenAI | 93.3% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 13 | GPT-5.4 (xhigh) | OpenAI | 93.3% | $2.50 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 14 | Grok 4.6 (xhigh) | xAI | 93.2% | $1.25 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 15 | Kimi K3 (max) | Moonshot | 93.1% | $1.29 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 16 | Ornith-1.5-397B | ornith-ai | 92.8% | — | HuggingFace GPQA Diamond leaderboard | 2026-10-05 |
| 17 | Grok 4.7 (xhigh) | xAI | 92.7% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 18 | Qwen3.8 Max (xhigh) | Alibaba | 92.7% | $1.65 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 19 | Gemini 3 Pro Preview | Google DeepMind | 92.6% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 20 | Qwen3.8-2.4T-A95B | Qwen | 92.6% | $2.00 | HuggingFace GPQA Diamond leaderboard | 2026-10-05 |
| 21 | Hy4-preview | tencent | 92.3% | $0.751 | HuggingFace GPQA Diamond leaderboard | 2026-10-05 |
| 22 | Qwen3.8 Max (0902) (xhigh) | Alibaba | 92.3% | $1.65 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 23 | Qwen3.8-Flash-Next | Qwen | 91.7% | — | HuggingFace GPQA Diamond leaderboard | 2026-10-05 |
| 24 | DeepSeek V4 Pro 0813 (max) | DeepSeek | 91.7% | $0.660 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 25 | GPT-5.6 Luna (max) | OpenAI | 91.6% | $0.200 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 26 | GPT-5.2 (xhigh) | OpenAI | 91.4% | $1.75 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 27 | DeepSeek V4 Flash 0731 (max) | DeepSeek | 91.0% | $0.015 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 28 | GLM-5.3 (max) | Z.ai (Zhipu AI) | 90.9% | $0.070 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 29 | MiniMax-M3 | MiniMax | 90.9% | $0.230 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 30 | Qwen3.7 Max | Alibaba | 90.9% | $1.25 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 31 | Darwin-398B-JGOS | FINAL-Bench | 90.9% | — | HuggingFace GPQA Diamond leaderboard | 2026-10-05 |
| 32 | Claude Opus 5.5 (max) | Anthropic | 90.6% | $4.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 33 | Kimi K2.6 | Moonshot | 90.5% | $0.650 | HuggingFace GPQA Diamond leaderboard | 2026-10-05 |
| 34 | GPT-6 Luna (max) | OpenAI | 90.5% | $0.100 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 35 | Hy3 | tencent | 90.4% | $0.083 | HuggingFace GPQA Diamond leaderboard | 2026-10-05 |
| 36 | GLM-5.3-Flash (max) | Z.ai (Zhipu AI) | 90.2% | $0.110 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 37 | Muse Spark | Meta AI | 89.8% | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 38 | DeepSeek v4 Pro (max) | DeepSeek | 89.6% | $0.435 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 39 | grok-4.20-0309-reasoning | xAI | 89.3% | $1.25 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 40 | Ornith-1.5-35B-A3B | ornith-ai | 89.2% | — | HuggingFace GPQA Diamond leaderboard | 2026-10-05 |
| 41 | grok-4.3 (high) | xAI | 88.8% | $1.25 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 42 | Inkling Small (xhigh) | Thinking Machines | 88.5% | $0.450 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 43 | Darwin-36B-Opus | FINAL-Bench | 88.4% | — | HuggingFace GPQA Diamond leaderboard | 2026-10-05 |
| 44 | Qwen3.5 397B-A17B | Alibaba | 88.4% | $0.164 | HuggingFace GPQA Diamond leaderboard | 2026-10-05 |
| 45 | Claude Opus 4.6 (max) | Anthropic | 88.4% | $5.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 46 | Qwen 3.6 Plus (2026-04-02) | Alibaba | 88.4% | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 47 | Ring-2.6-1T | inclusionAI | 88.3% | — | HuggingFace GPQA Diamond leaderboard | 2026-10-05 |
| 48 | Inkling (xhigh) | Thinking Machines | 88.3% | $0.950 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 49 | NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 | nvidia | 87.9% | — | HuggingFace GPQA Diamond leaderboard | 2026-10-05 |
| 50 | Kimi K2.7 Code | Moonshot | 87.9% | $0.671 | Epoch AI Benchmarking Hub | 2026-10-05 |
Compare all benchmarks side by side on the live leaderboard →
Among the ten highest scorers, the cheapest listed API price belongs to Gemini 3.8 Flash (high) at $0.750 per million input tokens (score 95.4%).
What GPQA Diamond measures
GPQA Diamond is the hardest subset of GPQA: 198 multiple-choice questions in biology, physics and chemistry written and checked by domain experts. The questions are designed to be "Google-proof": skilled non-experts with web access still do badly, so a high score points to real scientific reasoning rather than lookup.
How to read these scores
Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.
Where the data comes from
Epoch AI Benchmarking Hub (CC BY 4.0) and HuggingFace GPQA Diamond leaderboard. Scores were last captured 2026-10-05; the Captured column gives each row's date.