SciCode leaderboard: which LLMs score highest

SciCode (scientific research code) currently has scores for 100 models. The top score is 66.9% by Claude Opus 5.5 (max); the median model scores 43.2% and the lowest scores 10.8%.

SciCode leaderboard: top 50 of 100 models (highest-effort row per model)
#ModelOrganisationSciCode scoreCheapest input $/MSourceCaptured
1Claude Opus 5.5 (max)Anthropic66.9%$4.00Epoch AI Benchmarking Hub2026-10-05
2Claude Fable 5.1 (max)Anthropic63.1%$10.00Epoch AI Benchmarking Hub2026-10-05
3Claude Fable 5 (max)Anthropic61.0%$10.00Epoch AI Benchmarking Hub2026-10-05
4Claude Sonnet 5.5 (max)Anthropic61.0%$2.00Epoch AI Benchmarking Hub2026-10-05
5Kimi K3 (max)Moonshot59.5%$1.29Epoch AI Benchmarking Hub2026-10-05
6GLM-5.3 (max)Z.ai (Zhipu AI)59.0%$0.070Epoch AI Benchmarking Hub2026-10-05
7Muse Spark 1.3 (max)Meta AI58.8%$1.25Epoch AI Benchmarking Hub2026-10-05
8GPT-6 Sol (max)OpenAI57.6%$2.00Epoch AI Benchmarking Hub2026-10-05
9Grok 4.7 (xhigh)xAI57.4%$2.00Epoch AI Benchmarking Hub2026-10-05
10GPT-5.6 Sol (max)OpenAI57.1%$2.00Epoch AI Benchmarking Hub2026-10-05
11Gemini 3.7 Flash (high)Google DeepMind56.8%$0.750Epoch AI Benchmarking Hub2026-10-05
12Gemini 3.8 Flash (high)Google DeepMind56.6%$0.750Epoch AI Benchmarking Hub2026-10-05
13GPT-5.4 (xhigh)OpenAI56.6%$2.50Epoch AI Benchmarking Hub2026-10-05
14GPT-6 Astra (max)OpenAI56.5%$10.00Epoch AI Benchmarking Hub2026-10-05
15Claude Opus 5 (max)Anthropic56.4%$5.00Epoch AI Benchmarking Hub2026-10-05
16Muse Spark 1.2 (xhigh)Meta AI56.4%$1.25Epoch AI Benchmarking Hub2026-10-05
17GPT-5.5 (xhigh)OpenAI56.1%$5.00Epoch AI Benchmarking Hub2026-10-05
18GPT-5.6 Terra (max)OpenAI55.0%$2.00Epoch AI Benchmarking Hub2026-10-05
19GPT-6 Luna (max)OpenAI54.6%$0.100Epoch AI Benchmarking Hub2026-10-05
20Claude Opus 4.7 (max)Anthropic54.5%$5.00Epoch AI Benchmarking Hub2026-10-05
21GPT-6.1 Sol (max)OpenAI54.2%$2.00Epoch AI Benchmarking Hub2026-10-05
22Grok 4.5 (high)xAI54.1%$2.00Epoch AI Benchmarking Hub2026-10-05
23Claude Sonnet 5 (max)Anthropic53.6%$2.00Epoch AI Benchmarking Hub2026-10-05
24GPT-5.6 Luna (max)OpenAI53.6%$0.200Epoch AI Benchmarking Hub2026-10-05
25Kimi K2.6Moonshot53.5%$0.650Epoch AI Benchmarking Hub2026-10-05
26deepseek-v4.1-flash-maxDeepSeek51.9%—Epoch AI Benchmarking Hub2026-10-05
27Grok 4.6 (xhigh)xAI51.6%$1.25Epoch AI Benchmarking Hub2026-10-05
28Muse SparkMeta AI51.5%—Epoch AI Benchmarking Hub2026-10-05
29DeepSeek V4 Pro 0813 (max)DeepSeek51.0%$0.660Epoch AI Benchmarking Hub2026-10-05
30grok-build-0.1xAI50.2%$1.00Epoch AI Benchmarking Hub2026-10-05
31mimo-v2.5-proXiaomi Corp50.2%$0.435Epoch AI Benchmarking Hub2026-10-05
32DeepSeek v4 Pro (max)DeepSeek50.0%$0.435Epoch AI Benchmarking Hub2026-10-05
33DeepSeek V4 Flash 0731 (max)DeepSeek49.9%$0.015Epoch AI Benchmarking Hub2026-10-05
34GPT-5.4 mini (xhigh)OpenAI49.9%$0.750Epoch AI Benchmarking Hub2026-10-05
35Kimi K2.5Moonshot49.0%$0.450Epoch AI Benchmarking Hub2026-10-05
36Qwen3.7 MaxAlibaba48.8%$1.25Epoch AI Benchmarking Hub2026-10-05
37GPT-5.5 InstantOpenAI48.6%—Epoch AI Benchmarking Hub2026-10-05
38Kimi K2.7 CodeMoonshot47.5%$0.671Epoch AI Benchmarking Hub2026-10-05
39grok-4.3 (high)xAI47.3%$1.25Epoch AI Benchmarking Hub2026-10-05
40MiniMax-M3MiniMax47.1%$0.230Epoch AI Benchmarking Hub2026-10-05
41Inkling (xhigh)Thinking Machines47.0%$0.950Epoch AI Benchmarking Hub2026-10-05
42MiniMax-M2.7MiniMax47.0%$0.210Epoch AI Benchmarking Hub2026-10-05
43GPT-5.4 nano (xhigh)OpenAI46.9%$0.200Epoch AI Benchmarking Hub2026-10-05
44Claude Sonnet 4.6 (max)Anthropic46.8%$3.00Epoch AI Benchmarking Hub2026-10-05
45Qwen3.8-27B (XHigh)Alibaba46.6%$0.400Epoch AI Benchmarking Hub2026-10-05
46Qwen3.7 PlusAlibaba45.5%$0.282Epoch AI Benchmarking Hub2026-10-05
47GLM-4.7Z.ai (Zhipu AI)45.1%$0.400Epoch AI Benchmarking Hub2026-10-05
48DeepSeek v4 Flash (max)DeepSeek44.9%$0.090Epoch AI Benchmarking Hub2026-10-05
49GLM-5.1Z.ai (Zhipu AI)43.8%$1.05Epoch AI Benchmarking Hub2026-10-05
50Gemma 4 31B ITGoogle DeepMind43.4%$0.100Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to GLM-5.3 (max) at $0.070 per million input tokens (score 59.0%).

What SciCode measures

SciCode measures scientific research code.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks