LLM benchmarks: every leaderboard we track

An LLM benchmark is a fixed test that scores language models on the same tasks. LLMs Tiger lists 53 benchmarks that each have at least 15 models, from graduate-level science questions to real GitHub bug fixes. Pick one to see the full ranking, or read our guide to what these tests measure.

LLM benchmarks tracked by LLMs Tiger
BenchmarkWhat it measuresModelsLeaderTop score
GPQA Diamondexpert-level science questions274GPT-6 Astra (max)95.8%
Epoch ECIEpoch Capabilities Index: Epoch's own aggregate of ~40 benchmarks205Qwen 3.8 Max157
OTIS Mock AIMEcompetition maths (AIME 2024-25 style)188Claude Fable 5.1 (max)100.0%
DTBenchdynamic tool-use benchmark149Claude Opus 5.5 (max)98.9%
MMLU-Pro57-subject knowledge + reasoning136Darwin-180B-RSI88.1%
Chess Puzzleschess tactics puzzles124GPT-6 Astra (max)72.0%
WeirdMLunusual machine-learning tasks112GPT-6 Astra (pro, max)93.6%
CritPtunpublished research-level physics problems111GPT-5.6 Sol (max)32.3%
LMCAlanguage-model capability assessment111Claude Opus 5.5 (max)68.2%
SciCodescientific research code100Claude Opus 5.5 (max)66.9%
Aider Code EditingAider official benchmark92Claude 3.5-Sonnet-2024102284.2%
ARC-AGI-2abstract reasoning, ARC Prize public eval set89GPT-6 Astra (max)95.0%
ALE-BenchAtCoder heuristic-optimisation contests (performance rating)78GPT-6 Astra (max)2,951
SWE-bench Verifiedreal GitHub issues resolved73Claude Opus 4.7 (max)83.5%
FrontierMath T1-3research-level maths, tiers 1-369GPT-6 Astra (max)93.7%
SimpleQA Verifiedshort factual questions66GPT-6 Astra (max)75.6%
SimpleBenchtrick questions that test common sense64Claude Fable 5 (max)81.9%
WebDev Arenabuilding web apps, crowd-voted (Elo)63Claude Opus 5.5 (max)1,820
Mystery Gamesdeductive reasoning puzzles57GPT-6 Astra (max)84.0%
FrontierMath T4research-level maths, tier 452GPT-6.1 Sol (max)100.0%
Aider Polyglotedits across 6 languages (225 exercises)49gemini-2.5-pro-preview-06-0583.1%
Fiction.liveBenchlong-context story comprehension (16k)47grok-4-070994.4%
ProofBenchwriting mathematical proofs43Claude Fable 5.1 (max)100.0%
METR time horizonlength of task (minutes) done at 50% success34Claude Mythos Preview (Early)1,045 min
BALROGlong-horizon game-playing agents34GPT-6 Astra (max)68.3%
Humanity's Last Examexpert-level questions across domains29Muse Spark40.6%
LMArena Code Elocrowd-voted coding answers29Claude Opus 5.5 (max)1,815
LMArena Text Elohead-to-head chat preference (crowd votes)28gemini-4-argon-high1,525
Terminal-Bench 2.0real terminal tasks (agent + model)28Gemini 3 Pro Preview69.4%
LMArena Search Elocrowd-voted web-search answers28claude-opus-4-6-search1,253
EnigmaEvalmulti-step puzzle hunts26GPT-5 Pro18.8%
Furniture Assemblyplanning a physical assembly25Claude Opus 5.5 (max)83.3%
GDP.pdfeconomically valuable document tasks22GPT-5.6 Sol (max)30.7%
LMArena Vision Elocrowd-voted image understanding22Gemini 3.7 Flash (high)1,296
SWE-bench Litereal GitHub issues resolved20Claude 4 Sonnet56.7%
DeepSWElong-horizon software engineering19Gemini 3.8 Flash (high)73.8%
EBR-Benchevidence-based reasoning19GPT-6 Astra (max)76.2%
GSO-Benchsoftware performance optimisation19GPT-5.5 (xhigh)40.2%
Vending-Bench 2running a vending business for a year (final money, $)19Kimi K2.6$6,205
LMArena Document Elocrowd-voted questions about documents18Claude Fable 5.1 (max)1,513
LiveBench Agentic CodingLiveBench release 2026_06_2517claude-opus-4-8-xhigh-effort56.1%
LiveBench CodingLiveBench release 2026_06_2517GPT-5.2 Codex83.6%
LiveBench Data AnalysisLiveBench release 2026_06_2517GPT-5.5 (xhigh)81.6%
LiveBench IFLiveBench release 2026_06_2517gemini-3.1-pro-preview-high79.1%
LiveBench LanguageLiveBench release 2026_06_2517GPT-5.5 (xhigh)87.4%
LiveBench MathematicsLiveBench release 2026_06_2517GPT-5.5 (xhigh)95.9%
LiveBench ReasoningLiveBench release 2026_06_2517claude-opus-4-8-xhigh-effort89.7%
GSM8Kgrade-school math word problems17mimo-v2.5-pro99.6%
CL-benchlearning new rules from context16GPT-5.4 (xhigh)27.9%
FrontierSWEhard multi-hour software engineering16GPT-6 Astra (max)65.5%
Cybenchcybersecurity capture-the-flag tasks16grok-4-070943.0%
FrontierCodehard programming tasks15Claude Opus 5 (max)53.4%
Terminal-Bench 4.0real terminal tasks, agent+model config15Codex + GPT-6 Astra58.2%

Also tracked, too few models for a leaderboard yet

Aider Refactoring, APEX-Agents, Blueprint-Bench 2, CL-bench Life, CursorBench, DeepResearch Bench, ExploitBench, GBA-Eval, HellaSwag, MCP Atlas, MirrorCode, OSWorld, OSWorld 2, PostTrainBench, Remote Labor Index, Surface Evolver, SWE-bench Multilingual, SWE-bench Multimodal, SWE-bench Test, SWE-rebench, TheAgentCompany and WeirdML v3.