LMArena Vision Elo leaderboard: which LLMs score highest

LMArena Vision Elo (crowd-voted image understanding) currently has scores for 22 models. The top score is 1,296 by Gemini 3.7 Flash (high); the median model scores 1,279 and the lowest scores 1,251.

LMArena Vision Elo leaderboard: top 22 of 22 models (highest-effort row per model)
#ModelOrganisationLMArena Vision Elo scoreCheapest input $/MSourceCaptured
1Gemini 3.7 Flash (high)Google DeepMind1,296$0.750LMArena2026-10-05
2Muse SparkMeta AI1,294—LMArena2026-10-05
3Muse Spark 1.2 (xhigh)Meta AI1,293$1.25LMArena2026-10-05
4GPT-6.1 Sol (max)OpenAI1,291$2.00LMArena2026-10-05
5Gemini 3.8 Flash (high)Google DeepMind1,290$0.750LMArena2026-10-05
6Muse Spark 1.3 (max)Meta AI1,290$1.25LMArena2026-10-05
7Gemini 3 ProGoogle DeepMind1,289$2.00LMArena2026-10-05
8Claude Fable 5.1 (max)Anthropic1,288$10.00LMArena2026-10-05
9gemini-3.5-flash-highGoogle1,284$1.50LMArena2026-10-05
10GPT-6 Astra (max)OpenAI1,284$10.00LMArena2026-10-05
11Gemini 3.6 Flash (High)Google1,280$0.750LMArena2026-10-05
12gpt-5.2-chat-latest-20260210OpenAI1,278—LMArena2026-10-05
13GPT-5.5 InstantOpenAI1,276—LMArena2026-10-05
14Gemini 3 FlashGoogle DeepMind1,273$0.500LMArena2026-10-05
15Kimi K2.6Moonshot1,265$0.650LMArena2026-10-05
16Qwen3.7 PlusAlibaba1,265$0.282LMArena2026-10-05
17gemma-4-31bGoogle1,262$0.090LMArena2026-10-05
18gemini-3-flash (thinking-minimal)Google1,260$0.500LMArena2026-10-05
19dola-seed-2.0-pro—1,257—LMArena2026-10-05
20grok-4.20-beta-0309-reasoningxAI1,255$1.25LMArena2026-10-05
21kimi-k2.5-thinkingMoonshot AI1,252$0.450LMArena2026-10-05
22GPT-5.1 (high)OpenAI1,251$1.25LMArena2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to Gemini 3.7 Flash (high) at $0.750 per million input tokens (score 1,296).

What LMArena Vision Elo measures

LMArena Vision Elo measures crowd-voted image understanding.

How to read these scores

Scores are Elo-style ratings (higher is better). Only differences between models mean anything, and small gaps may sit inside the source's own margin of error. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

LMArena (third-party snapshot 2026-10-05) and LMArena (third-party snapshot 2026-10-04). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks