GSO-Bench leaderboard: which LLMs score highest
GSO-Bench (software performance optimisation) currently has scores for 19 models. The top score is 40.2% by GPT-5.5 (xhigh); the median model scores 4.9% and the lowest scores 0.0%.
| # | Model | Organisation | GSO-Bench score | Cheapest input $/M | Source | Captured |
|---|---|---|---|---|---|---|
| 1 | GPT-5.5 (xhigh) | OpenAI | 40.2% | $5.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 2 | GPT-5.4 (xhigh) | OpenAI | 31.4% | $2.50 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 3 | Gemini 3 Pro Preview | Google DeepMind | 18.6% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 4 | GPT-5.1 (high) | OpenAI | 13.7% | $1.25 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 5 | o3 (high) | OpenAI | 8.8% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 6 | Claude Opus 4 | Anthropic | 6.9% | $15.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 7 | GPT-5 (high) | OpenAI | 6.9% | $1.25 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 8 | Claude Sonnet 4 (unknown thinking) | Anthropic | 4.9% | $3.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 9 | claude-sonnet-4-20250514 | Anthropic | 4.9% | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 10 | Kimi K2 Instruct | Moonshot | 4.9% | $0.500 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 11 | Qwen3-Coder-480B-A35B-Instruct | Alibaba | 4.9% | $0.380 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 12 | Claude 3.5 Sonnet (Oct 2024) | Anthropic | 4.6% | $3.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 13 | Gemini 2.5 Pro Preview (Jun 2025) | Google DeepMind | 3.9% | $1.25 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 14 | claude-3-7-sonnet-20250219 | Anthropic | 3.8% | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 15 | o4-mini (high) | OpenAI | 3.6% | $1.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 16 | GLM-4.5-Air | Z.ai (Zhipu AI),Tsinghua University | 2.9% | $0.125 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 17 | o3-mini (high) | OpenAI | 1.3% | $1.10 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 18 | o3-mini-2025-01-31 (low) | OpenAI | 1.3% | $1.10 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 19 | GPT-4o (Nov 2024) | OpenAI | 0.0% | $2.50 | Epoch AI Benchmarking Hub | 2026-10-05 |
Compare all benchmarks side by side on the live leaderboard →
Among the ten highest scorers, the cheapest listed API price belongs to Kimi K2 Instruct at $0.500 per million input tokens (score 4.9%).
What GSO-Bench measures
GSO-Bench measures software performance optimisation.
How to read these scores
Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.
Where the data comes from
Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.