LLM benchmarks: every leaderboard we track
An LLM benchmark is a fixed test that scores language models on the same tasks. LLMs Tiger lists 53 benchmarks that each have at least 15 models, from graduate-level science questions to real GitHub bug fixes. Pick one to see the full ranking, or read our guide to what these tests measure.
| Benchmark | What it measures | Models | Leader | Top score |
|---|---|---|---|---|
| GPQA Diamond | expert-level science questions | 274 | GPT-6 Astra (max) | 95.8% |
| Epoch ECI | Epoch Capabilities Index: Epoch's own aggregate of ~40 benchmarks | 205 | Qwen 3.8 Max | 157 |
| OTIS Mock AIME | competition maths (AIME 2024-25 style) | 188 | Claude Fable 5.1 (max) | 100.0% |
| DTBench | dynamic tool-use benchmark | 149 | Claude Opus 5.5 (max) | 98.9% |
| MMLU-Pro | 57-subject knowledge + reasoning | 136 | Darwin-180B-RSI | 88.1% |
| Chess Puzzles | chess tactics puzzles | 124 | GPT-6 Astra (max) | 72.0% |
| WeirdML | unusual machine-learning tasks | 112 | GPT-6 Astra (pro, max) | 93.6% |
| CritPt | unpublished research-level physics problems | 111 | GPT-5.6 Sol (max) | 32.3% |
| LMCA | language-model capability assessment | 111 | Claude Opus 5.5 (max) | 68.2% |
| SciCode | scientific research code | 100 | Claude Opus 5.5 (max) | 66.9% |
| Aider Code Editing | Aider official benchmark | 92 | Claude 3.5-Sonnet-20241022 | 84.2% |
| ARC-AGI-2 | abstract reasoning, ARC Prize public eval set | 89 | GPT-6 Astra (max) | 95.0% |
| ALE-Bench | AtCoder heuristic-optimisation contests (performance rating) | 78 | GPT-6 Astra (max) | 2,951 |
| SWE-bench Verified | real GitHub issues resolved | 73 | Claude Opus 4.7 (max) | 83.5% |
| FrontierMath T1-3 | research-level maths, tiers 1-3 | 69 | GPT-6 Astra (max) | 93.7% |
| SimpleQA Verified | short factual questions | 66 | GPT-6 Astra (max) | 75.6% |
| SimpleBench | trick questions that test common sense | 64 | Claude Fable 5 (max) | 81.9% |
| WebDev Arena | building web apps, crowd-voted (Elo) | 63 | Claude Opus 5.5 (max) | 1,820 |
| Mystery Games | deductive reasoning puzzles | 57 | GPT-6 Astra (max) | 84.0% |
| FrontierMath T4 | research-level maths, tier 4 | 52 | GPT-6.1 Sol (max) | 100.0% |
| Aider Polyglot | edits across 6 languages (225 exercises) | 49 | gemini-2.5-pro-preview-06-05 | 83.1% |
| Fiction.liveBench | long-context story comprehension (16k) | 47 | grok-4-0709 | 94.4% |
| ProofBench | writing mathematical proofs | 43 | Claude Fable 5.1 (max) | 100.0% |
| METR time horizon | length of task (minutes) done at 50% success | 34 | Claude Mythos Preview (Early) | 1,045 min |
| BALROG | long-horizon game-playing agents | 34 | GPT-6 Astra (max) | 68.3% |
| Humanity's Last Exam | expert-level questions across domains | 29 | Muse Spark | 40.6% |
| LMArena Code Elo | crowd-voted coding answers | 29 | Claude Opus 5.5 (max) | 1,815 |
| LMArena Text Elo | head-to-head chat preference (crowd votes) | 28 | gemini-4-argon-high | 1,525 |
| Terminal-Bench 2.0 | real terminal tasks (agent + model) | 28 | Gemini 3 Pro Preview | 69.4% |
| LMArena Search Elo | crowd-voted web-search answers | 28 | claude-opus-4-6-search | 1,253 |
| EnigmaEval | multi-step puzzle hunts | 26 | GPT-5 Pro | 18.8% |
| Furniture Assembly | planning a physical assembly | 25 | Claude Opus 5.5 (max) | 83.3% |
| GDP.pdf | economically valuable document tasks | 22 | GPT-5.6 Sol (max) | 30.7% |
| LMArena Vision Elo | crowd-voted image understanding | 22 | Gemini 3.7 Flash (high) | 1,296 |
| SWE-bench Lite | real GitHub issues resolved | 20 | Claude 4 Sonnet | 56.7% |
| DeepSWE | long-horizon software engineering | 19 | Gemini 3.8 Flash (high) | 73.8% |
| EBR-Bench | evidence-based reasoning | 19 | GPT-6 Astra (max) | 76.2% |
| GSO-Bench | software performance optimisation | 19 | GPT-5.5 (xhigh) | 40.2% |
| Vending-Bench 2 | running a vending business for a year (final money, $) | 19 | Kimi K2.6 | $6,205 |
| LMArena Document Elo | crowd-voted questions about documents | 18 | Claude Fable 5.1 (max) | 1,513 |
| LiveBench Agentic Coding | LiveBench release 2026_06_25 | 17 | claude-opus-4-8-xhigh-effort | 56.1% |
| LiveBench Coding | LiveBench release 2026_06_25 | 17 | GPT-5.2 Codex | 83.6% |
| LiveBench Data Analysis | LiveBench release 2026_06_25 | 17 | GPT-5.5 (xhigh) | 81.6% |
| LiveBench IF | LiveBench release 2026_06_25 | 17 | gemini-3.1-pro-preview-high | 79.1% |
| LiveBench Language | LiveBench release 2026_06_25 | 17 | GPT-5.5 (xhigh) | 87.4% |
| LiveBench Mathematics | LiveBench release 2026_06_25 | 17 | GPT-5.5 (xhigh) | 95.9% |
| LiveBench Reasoning | LiveBench release 2026_06_25 | 17 | claude-opus-4-8-xhigh-effort | 89.7% |
| GSM8K | grade-school math word problems | 17 | mimo-v2.5-pro | 99.6% |
| CL-bench | learning new rules from context | 16 | GPT-5.4 (xhigh) | 27.9% |
| FrontierSWE | hard multi-hour software engineering | 16 | GPT-6 Astra (max) | 65.5% |
| Cybench | cybersecurity capture-the-flag tasks | 16 | grok-4-0709 | 43.0% |
| FrontierCode | hard programming tasks | 15 | Claude Opus 5 (max) | 53.4% |
| Terminal-Bench 4.0 | real terminal tasks, agent+model config | 15 | Codex + GPT-6 Astra | 58.2% |
Also tracked, too few models for a leaderboard yet
Aider Refactoring, APEX-Agents, Blueprint-Bench 2, CL-bench Life, CursorBench, DeepResearch Bench, ExploitBench, GBA-Eval, HellaSwag, MCP Atlas, MirrorCode, OSWorld, OSWorld 2, PostTrainBench, Remote Labor Index, Surface Evolver, SWE-bench Multilingual, SWE-bench Multimodal, SWE-bench Test, SWE-rebench, TheAgentCompany and WeirdML v3.