GSM8K leaderboard: which LLMs score highest

GSM8K (grade-school math word problems) currently has scores for 17 models. The top score is 99.6% by mimo-v2.5-pro; the median model scores 86.9% and the lowest scores 79.2%.

GSM8K leaderboard: top 17 of 17 models (highest-effort row per model)
#ModelOrganisationGSM8K scoreCheapest input $/MSourceCaptured
1mimo-v2.5-proXiaomi Corp99.6%$0.435HuggingFace GSM8K leaderboard2026-10-05
2Llama 3.1-405BMeta AI96.8%—HuggingFace GSM8K leaderboard2026-10-05
3granite-4.1-30bibm-granite94.2%—HuggingFace GSM8K leaderboard2026-10-05
4Llama-3.2-90B-Vision-InstructMeta AI93.1%$2.00HuggingFace GSM8K leaderboard2026-10-05
5granite-4.1-8bibm-granite92.5%$0.050HuggingFace GSM8K leaderboard2026-10-05
6Phi-3-medium-4k-instructmicrosoft91.0%$0.170HuggingFace GSM8K leaderboard2026-10-05
7Qwen2-72BAlibaba89.5%—HuggingFace GSM8K leaderboard2026-10-05
8DeepSeek-V3DeepSeek89.3%$0.200HuggingFace GSM8K leaderboard2026-10-05
9granite-4.1-3bibm-granite86.9%—HuggingFace GSM8K leaderboard2026-10-05
10Phi-3.5-mini-instructmicrosoft86.2%$0.130HuggingFace GSM8K leaderboard2026-10-05
11internlm2_5-7b-chatinternlm86.0%—HuggingFace GSM8K leaderboard2026-10-05
12Phi-3-mini-4k-instructMicrosoft85.7%$0.130HuggingFace GSM8K leaderboard2026-10-05
13Mellum2-12B-A2.5B-BaseJetBrains81.7%—HuggingFace GSM8K leaderboard2026-10-05
14Mellum2-12B-A2.5B-Base-PretrainJetBrains81.7%—HuggingFace GSM8K leaderboard2026-10-05
15Qwen2-7BQwen79.9%—HuggingFace GSM8K leaderboard2026-10-05
16internlm2-chat-20binternlm79.6%—HuggingFace GSM8K leaderboard2026-10-05
17DeepSeek-V2deepseek-ai79.2%—HuggingFace GSM8K leaderboard2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to granite-4.1-8b at $0.050 per million input tokens (score 92.5%).

What GSM8K measures

GSM8K is a set of grade-school maths word problems that need several arithmetic steps. Frontier models now score near the ceiling, so it mostly separates small and older models.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

HuggingFace GSM8K leaderboard. Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks