LMArena Document Elo leaderboard: which LLMs score highest
LMArena Document Elo (crowd-voted questions about documents) currently has scores for 18 models. The top score is 1,513 by Claude Fable 5.1 (max); the median model scores 1,440 and the lowest scores 1,403.
| # | Model | Organisation | LMArena Document Elo score | Cheapest input $/M | Source | Captured |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1 (max) | Anthropic | 1,513 | $10.00 | LMArena | 2026-10-05 |
| 2 | Muse Spark 1.3 (max) | Meta AI | 1,471 | $1.25 | LMArena | 2026-10-05 |
| 3 | GPT-6 Astra (max) | OpenAI | 1,468 | $10.00 | LMArena | 2026-10-05 |
| 4 | gemini-3.5-flash-high | 1,461 | $1.50 | LMArena | 2026-10-05 | |
| 5 | Gemini 3.6 Flash (High) | 1,456 | $0.750 | LMArena | 2026-10-05 | |
| 6 | Kimi K2.6 | Moonshot | 1,451 | $0.650 | LMArena | 2026-10-05 |
| 7 | claude-sonnet-4-5-20250929 | Anthropic | 1,450 | $3.00 | LMArena | 2026-10-05 |
| 8 | Muse Spark | Meta AI | 1,444 | — | LMArena | 2026-10-05 |
| 9 | Qwen3.7 Plus | Alibaba | 1,444 | $0.282 | LMArena | 2026-10-05 |
| 10 | MiniMax-M3 | MiniMax | 1,435 | $0.230 | LMArena | 2026-10-05 |
| 11 | Gemini 3 Pro | Google DeepMind | 1,434 | $2.00 | LMArena | 2026-10-05 |
| 12 | kimi-k2.5-thinking | Moonshot AI | 1,430 | $0.450 | LMArena | 2026-10-05 |
| 13 | gemma-4-31b | 1,425 | $0.090 | LMArena | 2026-10-05 | |
| 14 | claude-haiku-4-5-20251001 | Anthropic | 1,420 | $1.00 | LMArena | 2026-10-05 |
| 15 | glm-5v-turbo | Z.ai | 1,416 | $0.704 | LMArena | 2026-10-05 |
| 16 | grok-4.20-beta-0309-reasoning | xAI | 1,416 | $1.25 | LMArena | 2026-10-05 |
| 17 | Gemini 3 Flash | Google DeepMind | 1,413 | $0.500 | LMArena | 2026-10-05 |
| 18 | GPT-5.5 Instant | OpenAI | 1,403 | — | LMArena | 2026-10-05 |
Compare all benchmarks side by side on the live leaderboard →
Among the ten highest scorers, the cheapest listed API price belongs to MiniMax-M3 at $0.230 per million input tokens (score 1,435).
What LMArena Document Elo measures
LMArena Document Elo measures crowd-voted questions about documents.
How to read these scores
Scores are Elo-style ratings (higher is better). Only differences between models mean anything, and small gaps may sit inside the source's own margin of error. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.
Where the data comes from
LMArena (third-party snapshot 2026-10-05). Scores were last captured 2026-10-05; the Captured column gives each row's date.