SimpleBench leaderboard: which LLMs score highest

SimpleBench (trick questions that test common sense) currently has scores for 64 models. The top score is 81.9% by Claude Fable 5 (max); the median model scores 43.3% and the lowest scores 10.7%.

SimpleBench leaderboard: top 50 of 64 models (highest-effort row per model)
#ModelOrganisationSimpleBench scoreCheapest input $/MSourceCaptured
1Claude Fable 5 (max)Anthropic81.9%$10.00Epoch AI Benchmarking Hub2026-10-05
2Gemini 3 Pro PreviewGoogle DeepMind76.4%$2.00Epoch AI Benchmarking Hub2026-10-05
3GPT-5.6 Sol (pro, unknown thinking)OpenAI71.7%$2.00Epoch AI Benchmarking Hub2026-10-05
4GPT-5.6 Sol (pro, xhigh)OpenAI71.7%$2.00Epoch AI Benchmarking Hub2026-10-05
5Qwen3.7 MaxAlibaba70.4%$1.25Epoch AI Benchmarking Hub2026-10-05
6Qwen 3.6 Max (Preview)Alibaba63.0%—Epoch AI Benchmarking Hub2026-10-05
7Gemini 2.5 Pro Preview (Jun 2025)Google DeepMind62.4%$1.25Epoch AI Benchmarking Hub2026-10-05
8GPT-5 ProOpenAI61.6%$15.00Epoch AI Benchmarking Hub2026-10-05
9Kimi K3 (max)Moonshot60.7%$1.29Epoch AI Benchmarking Hub2026-10-05
10grok-4-0709xAI60.5%—Epoch AI Benchmarking Hub2026-10-05
11Claude Opus 4.1 (unknown thinking)Anthropic60.0%$15.00Epoch AI Benchmarking Hub2026-10-05
12claude-opus-4-1-20250805Anthropic60.0%—Epoch AI Benchmarking Hub2026-10-05
13Claude Opus 4Anthropic58.8%$15.00Epoch AI Benchmarking Hub2026-10-05
14Kimi K2.7 CodeMoonshot57.9%$0.671Epoch AI Benchmarking Hub2026-10-05
15GPT-5 (high)OpenAI56.7%$1.25Epoch AI Benchmarking Hub2026-10-05
16grok-4-1-fast-non-reasoningxAI56.0%$0.200Epoch AI Benchmarking Hub2026-10-05
17grok-4-1-fast-reasoningxAI56.0%$0.200Epoch AI Benchmarking Hub2026-10-05
18GLM-5.1Z.ai (Zhipu AI)55.1%$1.05Epoch AI Benchmarking Hub2026-10-05
19GLM-5Z.ai (Zhipu AI)53.2%$0.600Epoch AI Benchmarking Hub2026-10-05
20GPT-5.1 (high)OpenAI53.2%$1.25Epoch AI Benchmarking Hub2026-10-05
21o3 (high)OpenAI53.1%$2.00Epoch AI Benchmarking Hub2026-10-05
22DeepSeek-V3.2-SpecialeDeepSeek52.6%$0.580Epoch AI Benchmarking Hub2026-10-05
23Gemini 2.5 Pro Exp (Mar 2025)Google DeepMind51.6%—Epoch AI Benchmarking Hub2026-10-05
24Gemini 2.5 Pro Preview (Mar 2025)Google DeepMind51.6%$1.25Epoch AI Benchmarking Hub2026-10-05
25GLM-4.7Z.ai (Zhipu AI)47.7%$0.400Epoch AI Benchmarking Hub2026-10-05
26GLM-4.7 (Novita)Z.ai (Zhipu AI)47.7%$0.400Epoch AI Benchmarking Hub2026-10-05
27Kimi K2.5Moonshot46.8%$0.450Epoch AI Benchmarking Hub2026-10-05
28DeepSeek v4 (unknown)DeepSeek46.3%—Epoch AI Benchmarking Hub2026-10-05
29MiniMax-M3MiniMax45.8%$0.230Epoch AI Benchmarking Hub2026-10-05
30Claude Sonnet 4 (unknown thinking)Anthropic45.5%$3.00Epoch AI Benchmarking Hub2026-10-05
31claude-sonnet-4-20250514_12KAnthropic45.5%—Epoch AI Benchmarking Hub2026-10-05
32claude-3-7-sonnet-20250219Anthropic44.9%—Epoch AI Benchmarking Hub2026-10-05
33o1-previewOpenAI41.7%$15.00Epoch AI Benchmarking Hub2026-10-05
34Claude 3.5 Sonnet (Oct 2024)Anthropic41.4%$3.00Epoch AI Benchmarking Hub2026-10-05
35gemini-2.5-flashGoogle DeepMind41.2%$0.300Epoch AI Benchmarking Hub2026-10-05
36DeepSeek-R1 (May 2025)DeepSeek40.8%$0.400Epoch AI Benchmarking Hub2026-10-05
37o1 (high)OpenAI40.1%$15.00Epoch AI Benchmarking Hub2026-10-05
38DeepSeek-V3.1DeepSeek40.0%$0.250Epoch AI Benchmarking Hub2026-10-05
39o4-mini (high)OpenAI38.7%$1.00Epoch AI Benchmarking Hub2026-10-05
40grok-3xAI36.1%$3.00Epoch AI Benchmarking Hub2026-10-05
41Qwen 3.6 Flash (2026-04-16)Alibaba35.2%—Epoch AI Benchmarking Hub2026-10-05
42GPT-4.5 Preview (Feb 2025)OpenAI34.5%$75.00Epoch AI Benchmarking Hub2026-10-05
43gemini-exp-1206Google DeepMind,Google31.1%$0.300Epoch AI Benchmarking Hub2026-10-05
44qwen3-235b-a22bAlibaba31.0%$0.180Epoch AI Benchmarking Hub2026-10-05
45DeepSeek-R1DeepSeek30.9%$0.400Epoch AI Benchmarking Hub2026-10-05
46Gemini 2.0 Flash Thinking ExpGoogle DeepMind,Google30.7%—Epoch AI Benchmarking Hub2026-10-05
47Llama 4 MaverickMeta AI27.7%$0.188Epoch AI Benchmarking Hub2026-10-05
48Claude 3.5 Sonnet (Jun 2024)Anthropic27.5%$3.00Epoch AI Benchmarking Hub2026-10-05
49DeepSeek-V3 (Mar 2025)DeepSeek27.2%$0.200Epoch AI Benchmarking Hub2026-10-05
50gemini-1.5-pro-002Google DeepMind27.1%—Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to Qwen3.7 Max at $1.25 per million input tokens (score 70.4%).

What SimpleBench measures

SimpleBench poses short questions with a trick or trap that most people catch using common sense. It is built to expose models that pattern-match instead of reason.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks