WebDev Arena leaderboard: which LLMs score highest

WebDev Arena (building web apps, crowd-voted (Elo)) currently has scores for 63 models. The top score is 1,820 by Claude Opus 5.5 (max); the median model scores 1,399 and the lowest scores 1,149.

WebDev Arena leaderboard: top 50 of 63 models (highest-effort row per model)
#ModelOrganisationWebDev Arena scoreCheapest input $/MSourceCaptured
1Claude Opus 5.5 (max)Anthropic1,820$4.00Epoch AI Benchmarking Hub2026-10-05
2GPT-6 Astra (max)OpenAI1,800$10.00Epoch AI Benchmarking Hub2026-10-05
3GPT-6.1 Sol (max)OpenAI1,759$2.00Epoch AI Benchmarking Hub2026-10-05
4Claude Fable 5.1 (max)Anthropic1,758$10.00Epoch AI Benchmarking Hub2026-10-05
5GPT-6 Sol (max)OpenAI1,689$2.00Epoch AI Benchmarking Hub2026-10-05
6Claude Opus 5 (max)Anthropic1,687$5.00Epoch AI Benchmarking Hub2026-10-05
7Kimi K3 (max)Moonshot1,674$1.29Epoch AI Benchmarking Hub2026-10-05
8Muse Spark 1.3 (max)Meta AI1,652$1.25Epoch AI Benchmarking Hub2026-10-05
9GLM-5.3 (max)Z.ai (Zhipu AI)1,614$0.070Epoch AI Benchmarking Hub2026-10-05
10deepseek-v4.1-flash-maxDeepSeek1,614—Epoch AI Benchmarking Hub2026-10-05
11Gemini 3.7 Flash (high)Google DeepMind1,587$0.750Epoch AI Benchmarking Hub2026-10-05
12GPT-6 Luna (max)OpenAI1,582$0.100Epoch AI Benchmarking Hub2026-10-05
13Gemini 3.8 Flash (high)Google DeepMind1,568$0.750Epoch AI Benchmarking Hub2026-10-05
14Muse Spark 1.2 (xhigh)Meta AI1,534$1.25Epoch AI Benchmarking Hub2026-10-05
15Qwen3.7 MaxAlibaba1,517$1.25Epoch AI Benchmarking Hub2026-10-05
16Muse SparkMeta AI1,513—Epoch AI Benchmarking Hub2026-10-05
17GPT-5.5 (xhigh)OpenAI1,510$5.00Epoch AI Benchmarking Hub2026-10-05
18Kimi K2.6Moonshot1,509$0.650Epoch AI Benchmarking Hub2026-10-05
19GLM-5.1Z.ai (Zhipu AI)1,508$1.05Epoch AI Benchmarking Hub2026-10-05
20MiniMax-M3MiniMax1,487$0.230Epoch AI Benchmarking Hub2026-10-05
21Qwen 3.6 Max (Preview)Alibaba1,479—Epoch AI Benchmarking Hub2026-10-05
22mimo-v2.5-proXiaomi Corp1,475$0.435Epoch AI Benchmarking Hub2026-10-05
23Kimi K2.7 CodeMoonshot1,473$0.671Epoch AI Benchmarking Hub2026-10-05
24Qwen 3.6 Plus (2026-04-02)Alibaba1,461—Epoch AI Benchmarking Hub2026-10-05
25Gemini 3 Pro PreviewGoogle DeepMind1,439$2.00Epoch AI Benchmarking Hub2026-10-05
26mimo-v2.5Xiaomi Corp1,437$0.140Epoch AI Benchmarking Hub2026-10-05
27Kimi K2.5Moonshot1,436$0.450Epoch AI Benchmarking Hub2026-10-05
28GLM-5Z.ai (Zhipu AI)1,436$0.600Epoch AI Benchmarking Hub2026-10-05
29GLM-4.7Z.ai (Zhipu AI)1,435$0.400Epoch AI Benchmarking Hub2026-10-05
30mimo-v2-proXiaomi Corp1,433$1.10Epoch AI Benchmarking Hub2026-10-05
31Kimi K2.5 (instant)Moonshot1,406$0.450Epoch AI Benchmarking Hub2026-10-05
32Qwen3.5 PlusAlibaba1,399—Epoch AI Benchmarking Hub2026-10-05
33MiniMax-M2.7MiniMax1,398$0.210Epoch AI Benchmarking Hub2026-10-05
34Claude Opus 4.1 (unknown thinking)Anthropic1,389$15.00Epoch AI Benchmarking Hub2026-10-05
35MiniMax-M2.1MiniMax1,387$0.300Epoch AI Benchmarking Hub2026-10-05
36claude-opus-4-1-20250805Anthropic1,386—Epoch AI Benchmarking Hub2026-10-05
37MiniMax-M2.5MiniMax1,384$0.270Epoch AI Benchmarking Hub2026-10-05
38grok-4-20xAI1,374$1.25Epoch AI Benchmarking Hub2026-10-05
39Gemma 4 31B ITGoogle DeepMind1,364$0.100Epoch AI Benchmarking Hub2026-10-05
40Gemma 4 26B A4BGoogle DeepMind1,361$0.076Epoch AI Benchmarking Hub2026-10-05
41Qwen3.5 27BAlibaba1,357$0.195Epoch AI Benchmarking Hub2026-10-05
42Laguna M.1Poolside1,347—Epoch AI Benchmarking Hub2026-10-05
43GLM-4.6Z.ai (Zhipu AI),Tsinghua University1,341$0.430Epoch AI Benchmarking Hub2026-10-05
44GPT-5.2 CodexOpenAI1,339$1.75Epoch AI Benchmarking Hub2026-10-05
45Kimi K2 Thinking TurboMoonshot1,337—Epoch AI Benchmarking Hub2026-10-05
46GPT-5.1 CodexOpenAI1,336$1.25Epoch AI Benchmarking Hub2026-10-05
47MiMo-V2-FlashXiaomiMiMo1,331$0.110Epoch AI Benchmarking Hub2026-10-05
48DeepSeek-V3.2 (Thinking; Novita)DeepSeek1,325$0.259Epoch AI Benchmarking Hub2026-10-05
49Laguna XS.2Poolside1,302—Epoch AI Benchmarking Hub2026-10-05
50MiniMax-M2MiniMax1,298$0.300Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to GLM-5.3 (max) at $0.070 per million input tokens (score 1,614).

What WebDev Arena measures

WebDev Arena asks two anonymous models to build the same web app and lets users vote on the better result; the ratings are Elo-style.

How to read these scores

Scores are Elo-style ratings (higher is better). Only differences between models mean anything, and small gaps may sit inside the source's own margin of error. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks