FrontierSWE leaderboard: which LLMs score highest

FrontierSWE (hard multi-hour software engineering) currently has scores for 16 models. The top score is 65.5% by GPT-6 Astra (max); the median model scores 29.8% and the lowest scores 4.1%.

FrontierSWE leaderboard: top 16 of 16 models (highest-effort row per model)
#ModelOrganisationFrontierSWE scoreCheapest input $/MSourceCaptured
1GPT-6 Astra (max)OpenAI65.5%$10.00Epoch AI Benchmarking Hub2026-10-05
2Claude Opus 5.5 (max)Anthropic62.3%$4.00Epoch AI Benchmarking Hub2026-10-05
3Claude Sonnet 5.5 (max)Anthropic61.9%$2.00Epoch AI Benchmarking Hub2026-10-05
4Claude Fable 5.1 (max)Anthropic56.3%$10.00Epoch AI Benchmarking Hub2026-10-05
5Claude Opus 5 (max)Anthropic52.0%$5.00Epoch AI Benchmarking Hub2026-10-05
6Claude Fable 5 (max)Anthropic47.0%$10.00Epoch AI Benchmarking Hub2026-10-05
7GPT-5.6 Sol (max)OpenAI32.2%$2.00Epoch AI Benchmarking Hub2026-10-05
8GLM-5.3 (max)Z.ai (Zhipu AI)30.2%$0.070Epoch AI Benchmarking Hub2026-10-05
9Grok 4.7 (xhigh)xAI29.5%$2.00Epoch AI Benchmarking Hub2026-10-05
10Kimi K3 (max)Moonshot25.9%$1.29Epoch AI Benchmarking Hub2026-10-05
11Grok 4.6 (xhigh)xAI25.3%$1.25Epoch AI Benchmarking Hub2026-10-05
12Gemini 3.7 Flash (high)Google DeepMind20.3%$0.750Epoch AI Benchmarking Hub2026-10-05
13Gemini 3.8 Flash (high)Google DeepMind19.6%$0.750Epoch AI Benchmarking Hub2026-10-05
14Qwen3.8 Max (xhigh)Alibaba15.8%$1.65Epoch AI Benchmarking Hub2026-10-05
15Muse Spark 1.2 (xhigh)Meta AI12.0%$1.25Epoch AI Benchmarking Hub2026-10-05
16Inkling (xhigh)Thinking Machines4.1%$0.950Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to GLM-5.3 (max) at $0.070 per million input tokens (score 30.2%).

What FrontierSWE measures

FrontierSWE measures hard multi-hour software engineering.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks