SimpleQA Verified leaderboard: which LLMs score highest

SimpleQA Verified (short factual questions) currently has scores for 66 models. The top score is 75.6% by GPT-6 Astra (max); the median model scores 41.0% and the lowest scores 6.0%.

SimpleQA Verified leaderboard: top 50 of 66 models (highest-effort row per model)
#ModelOrganisationSimpleQA Verified scoreCheapest input $/MSourceCaptured
1GPT-6 Astra (max)OpenAI75.6%$10.00Epoch AI Benchmarking Hub2026-10-05
2GPT-6.1 Sol (max)OpenAI73.9%$2.00Epoch AI Benchmarking Hub2026-10-05
3Claude Opus 5.5 (max)Anthropic72.2%$4.00Epoch AI Benchmarking Hub2026-10-05
4Claude Fable 5.1 (max)Anthropic70.8%$10.00Epoch AI Benchmarking Hub2026-10-05
5Gemini 3.8 Flash (high)Google DeepMind69.7%$0.750Epoch AI Benchmarking Hub2026-10-05
6GPT-5.6 Sol (max)OpenAI69.7%$2.00Epoch AI Benchmarking Hub2026-10-05
7Gemini 3.7 Flash (high)Google DeepMind69.2%$0.750Epoch AI Benchmarking Hub2026-10-05
8GPT-5.5 (xhigh)OpenAI63.0%$5.00Epoch AI Benchmarking Hub2026-10-05
9GPT-6 Sol (max)OpenAI60.7%$2.00Epoch AI Benchmarking Hub2026-10-05
10Muse Spark 1.2 (xhigh)Meta AI60.3%$1.25Epoch AI Benchmarking Hub2026-10-05
11Claude Opus 5 (max)Anthropic59.9%$5.00Epoch AI Benchmarking Hub2026-10-05
12Grok 4.7 (xhigh)xAI56.0%$2.00Epoch AI Benchmarking Hub2026-10-05
13Qwen3.7 MaxAlibaba55.8%$1.25Epoch AI Benchmarking Hub2026-10-05
14DeepSeek V4 Pro 0813 (max)DeepSeek52.9%$0.660Epoch AI Benchmarking Hub2026-10-05
15Qwen 3.6 Max (Preview)Alibaba52.0%—Epoch AI Benchmarking Hub2026-10-05
16Kimi K3 (max)Moonshot50.6%$1.29Epoch AI Benchmarking Hub2026-10-05
17GPT-5 (high)OpenAI50.1%$1.25Epoch AI Benchmarking Hub2026-10-05
18o3 (high)OpenAI49.4%$2.00Epoch AI Benchmarking Hub2026-10-05
19Grok 4.6 (xhigh)xAI48.9%$1.25Epoch AI Benchmarking Hub2026-10-05
20Qwen3-Max-InstructAlibaba48.7%—Epoch AI Benchmarking Hub2026-10-05
21Grok 4.5 (high)xAI48.3%$2.00Epoch AI Benchmarking Hub2026-10-05
22GPT-5.1 (high)OpenAI48.0%$1.25Epoch AI Benchmarking Hub2026-10-05
23Qwen3.8 Max (0902) (xhigh)Alibaba47.3%$1.65Epoch AI Benchmarking Hub2026-10-05
24Claude Opus 4.6 (max)Anthropic47.0%$5.00Epoch AI Benchmarking Hub2026-10-05
25DeepSeek v4 Pro (max)DeepSeek47.0%$0.435Epoch AI Benchmarking Hub2026-10-05
26Claude Sonnet 5.5 (max)Anthropic46.5%$2.00Epoch AI Benchmarking Hub2026-10-05
27GPT-5.4 Pro (xhigh)OpenAI46.3%$30.00Epoch AI Benchmarking Hub2026-10-05
28Qwen3.8 Max (xhigh)Alibaba45.8%$1.65Epoch AI Benchmarking Hub2026-10-05
29GPT-5.4 (xhigh)OpenAI45.1%$2.50Epoch AI Benchmarking Hub2026-10-05
30Qwen 3.6 Plus (2026-04-02)Alibaba44.1%—Epoch AI Benchmarking Hub2026-10-05
31GPT-5.6 Terra (max)OpenAI43.2%$2.00Epoch AI Benchmarking Hub2026-10-05
32GPT-6 Luna (max)OpenAI41.4%$0.100Epoch AI Benchmarking Hub2026-10-05
33o1 (high)OpenAI41.1%$15.00Epoch AI Benchmarking Hub2026-10-05
34GLM-5.3 (max)Z.ai (Zhipu AI)41.0%$0.070Epoch AI Benchmarking Hub2026-10-05
35GPT-5.6 Luna (max)OpenAI41.0%$0.200Epoch AI Benchmarking Hub2026-10-05
36Qwen3-235B-A22B-Thinking-2507Alibaba40.4%$0.230Epoch AI Benchmarking Hub2026-10-05
37Inkling (xhigh)Thinking Machines40.3%$0.950Epoch AI Benchmarking Hub2026-10-05
38GPT-5.2 (xhigh)OpenAI37.1%$1.75Epoch AI Benchmarking Hub2026-10-05
39Kimi K2.7 CodeMoonshot36.5%$0.671Epoch AI Benchmarking Hub2026-10-05
40Kimi K2.6Moonshot34.9%$0.650Epoch AI Benchmarking Hub2026-10-05
41Kimi K2.5Moonshot34.3%$0.450Epoch AI Benchmarking Hub2026-10-05
42GLM-5.1Z.ai (Zhipu AI)34.0%$1.05Epoch AI Benchmarking Hub2026-10-05
43Claude Sonnet 5 (max)Anthropic33.7%$2.00Epoch AI Benchmarking Hub2026-10-05
44DeepSeek V4 Flash 0731 (max)DeepSeek33.6%$0.015Epoch AI Benchmarking Hub2026-10-05
45grok-4.3 (high)xAI33.2%$1.25Epoch AI Benchmarking Hub2026-10-05
46Claude Sonnet 4.6 (max)Anthropic32.8%$3.00Epoch AI Benchmarking Hub2026-10-05
47GLM-4.7Z.ai (Zhipu AI)32.2%$0.400Epoch AI Benchmarking Hub2026-10-05
48GPT-4.1OpenAI31.1%$2.00Epoch AI Benchmarking Hub2026-10-05
49Claude Sonnet 4.5 (59k thinking)Anthropic30.7%$3.00Epoch AI Benchmarking Hub2026-10-05
50grok-4.20-0309-reasoningxAI30.2%$1.25Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to Gemini 3.8 Flash (high) at $0.750 per million input tokens (score 69.7%).

What SimpleQA Verified measures

SimpleQA Verified is a curated version of SimpleQA: short fact-seeking questions that each have one verifiable answer. It works as a rough check on how often a model states things that are not true.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks