GSO-Bench leaderboard: which LLMs score highest

GSO-Bench (software performance optimisation) currently has scores for 19 models. The top score is 40.2% by GPT-5.5 (xhigh); the median model scores 4.9% and the lowest scores 0.0%.

GSO-Bench leaderboard: top 19 of 19 models (highest-effort row per model)
#ModelOrganisationGSO-Bench scoreCheapest input $/MSourceCaptured
1GPT-5.5 (xhigh)OpenAI40.2%$5.00Epoch AI Benchmarking Hub2026-10-05
2GPT-5.4 (xhigh)OpenAI31.4%$2.50Epoch AI Benchmarking Hub2026-10-05
3Gemini 3 Pro PreviewGoogle DeepMind18.6%$2.00Epoch AI Benchmarking Hub2026-10-05
4GPT-5.1 (high)OpenAI13.7%$1.25Epoch AI Benchmarking Hub2026-10-05
5o3 (high)OpenAI8.8%$2.00Epoch AI Benchmarking Hub2026-10-05
6Claude Opus 4Anthropic6.9%$15.00Epoch AI Benchmarking Hub2026-10-05
7GPT-5 (high)OpenAI6.9%$1.25Epoch AI Benchmarking Hub2026-10-05
8Claude Sonnet 4 (unknown thinking)Anthropic4.9%$3.00Epoch AI Benchmarking Hub2026-10-05
9claude-sonnet-4-20250514Anthropic4.9%—Epoch AI Benchmarking Hub2026-10-05
10Kimi K2 InstructMoonshot4.9%$0.500Epoch AI Benchmarking Hub2026-10-05
11Qwen3-Coder-480B-A35B-InstructAlibaba4.9%$0.380Epoch AI Benchmarking Hub2026-10-05
12Claude 3.5 Sonnet (Oct 2024)Anthropic4.6%$3.00Epoch AI Benchmarking Hub2026-10-05
13Gemini 2.5 Pro Preview (Jun 2025)Google DeepMind3.9%$1.25Epoch AI Benchmarking Hub2026-10-05
14claude-3-7-sonnet-20250219Anthropic3.8%—Epoch AI Benchmarking Hub2026-10-05
15o4-mini (high)OpenAI3.6%$1.00Epoch AI Benchmarking Hub2026-10-05
16GLM-4.5-AirZ.ai (Zhipu AI),Tsinghua University2.9%$0.125Epoch AI Benchmarking Hub2026-10-05
17o3-mini (high)OpenAI1.3%$1.10Epoch AI Benchmarking Hub2026-10-05
18o3-mini-2025-01-31 (low)OpenAI1.3%$1.10Epoch AI Benchmarking Hub2026-10-05
19GPT-4o (Nov 2024)OpenAI0.0%$2.50Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to Kimi K2 Instruct at $0.500 per million input tokens (score 4.9%).

What GSO-Bench measures

GSO-Bench measures software performance optimisation.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks