Vending-Bench 2 leaderboard: which LLMs score highest

Vending-Bench 2 (running a vending business for a year (final money, $)) currently has scores for 19 models. The top score is $6,205 by Kimi K2.6; the median model scores $2,377 and the lowest scores $1.

Vending-Bench 2 leaderboard: top 19 of 19 models (highest-effort row per model)
#ModelOrganisationVending-Bench 2 scoreCheapest input $/MSourceCaptured
1Kimi K2.6Moonshot$6,205$0.650Epoch AI Benchmarking Hub2026-10-05
2GLM-5.1Z.ai (Zhipu AI)$5,634$1.05Epoch AI Benchmarking Hub2026-10-05
3Gemini 3 Pro PreviewGoogle DeepMind$5,478$2.00Epoch AI Benchmarking Hub2026-10-05
4Qwen 3.6 Plus (2026-04-02)Alibaba$5,115—Epoch AI Benchmarking Hub2026-10-05
5Kimi K2.7 CodeMoonshot$5,083$0.671Epoch AI Benchmarking Hub2026-10-05
6Claude Fable 5 (max)Anthropic$4,967$10.00Epoch AI Benchmarking Hub2026-10-05
7grok-4-20xAI$4,663$1.25Epoch AI Benchmarking Hub2026-10-05
8GLM-5Z.ai (Zhipu AI)$4,432$0.600Epoch AI Benchmarking Hub2026-10-05
9Qwen 3.6 Max (Preview)Alibaba$4,254—Epoch AI Benchmarking Hub2026-10-05
10GLM-4.7Z.ai (Zhipu AI)$2,377$0.400Epoch AI Benchmarking Hub2026-10-05
11MiniMax-M3MiniMax$2,158$0.230Epoch AI Benchmarking Hub2026-10-05
12Kimi K2.5Moonshot$1,198$0.450Epoch AI Benchmarking Hub2026-10-05
13grok-4-1-fast-reasoningxAI$1,107$0.200Epoch AI Benchmarking Hub2026-10-05
14gemini-2.5-flashGoogle DeepMind$549$0.300Epoch AI Benchmarking Hub2026-10-05
15Qwen3.5 FlashAlibaba$463$0.065Epoch AI Benchmarking Hub2026-10-05
16Qwen3.5 27BAlibaba$202$0.195Epoch AI Benchmarking Hub2026-10-05
17MiniMax-M2MiniMax$161$0.300Epoch AI Benchmarking Hub2026-10-05
18Qwen3-Max-InstructAlibaba$72—Epoch AI Benchmarking Hub2026-10-05
19Qwen3.5 PlusAlibaba$1—Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to GLM-4.7 at $0.400 per million input tokens (score $2,377).

What Vending-Bench 2 measures

Vending-Bench 2 measures running a vending business for a year (final money, $).

How to read these scores

Higher scores are better; the scale is specific to this benchmark, so compare models only on this board. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks