CursorBench leaderboard: which LLMs score highest

CursorBench (real-world coding-agent tasks) currently has scores for 15 models. The top score is 57.8% by Claude Opus 5.5 (max); the median model scores 41.6% and the lowest scores 27.7%.

CursorBench leaderboard: top 15 of 15 models (highest-effort row per model)
#ModelOrganisationCursorBench scoreCheapest input $/MSourceCaptured
1Claude Opus 5.5 (max)Anthropic57.8%$4.00Epoch AI Benchmarking Hub2026-10-06
2Claude Sonnet 5.5 (max)Anthropic55.5%$2.00Epoch AI Benchmarking Hub2026-10-06
3Claude Fable 5.1 (max)Anthropic51.8%$10.00Epoch AI Benchmarking Hub2026-10-06
4Claude Opus 5 (max)Anthropic46.6%$5.00Epoch AI Benchmarking Hub2026-10-06
5Grok 4.7 (xhigh)xAI46.3%$2.00Epoch AI Benchmarking Hub2026-10-06
6GLM-5.3 (max)Z.ai (Zhipu AI)42.6%$0.070Epoch AI Benchmarking Hub2026-10-06
7GPT-5.6 Sol (max)OpenAI41.7%$2.00Epoch AI Benchmarking Hub2026-10-06
8Muse Spark 1.3 (max)Meta AI41.6%$1.25Epoch AI Benchmarking Hub2026-10-06
9Grok 4.6 (xhigh)xAI41.4%$1.25Epoch AI Benchmarking Hub2026-10-06
10GPT-5.6 Terra (max)OpenAI41.3%$2.00Epoch AI Benchmarking Hub2026-10-06
11Gemini 3.8 Flash (high)Google DeepMind39.6%$0.750Epoch AI Benchmarking Hub2026-10-06
12GLM-5.3-Flash (max)Z.ai (Zhipu AI)36.8%$0.036Epoch AI Benchmarking Hub2026-10-06
13GPT-5.6 Luna (max)OpenAI35.9%$0.200Epoch AI Benchmarking Hub2026-10-06
14Claude Sonnet 5 (max)Anthropic34.1%$2.00Epoch AI Benchmarking Hub2026-10-06
15Composer 2.5Cursor27.7%—Epoch AI Benchmarking Hub2026-10-06

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to GLM-5.3 (max) at $0.070 per million input tokens (score 42.6%).

What CursorBench measures

CursorBench measures real-world coding-agent tasks.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-06; the Captured column gives each row's date.

Related benchmarks