LLM benchmarks explained: what they measure and how to read them
An LLM benchmark is a fixed set of tasks with a fixed way of grading, used to score large language models so they can be compared. Each model is given the same tasks and graded the same way, and the result is a number: a percentage correct, an Elo-style rating or a time. This guide covers the main kinds of benchmark, how to read their scores and why two sites can report different numbers for the same model.
The live numbers are on the LLMs Tiger leaderboard, which combines 75 public benchmarks.
What are the main types of LLM benchmark?
Knowledge and exam questions
These check how much a model knows and whether it can apply it. MMLU-Pro covers many academic subjects; GPQA Diamond uses graduate-level science questions; Humanity's Last Exam collects very hard expert-written questions. Older tests such as GSM8K are close to saturated, meaning the best models all score near the maximum, so they no longer separate leaders.
Reasoning and maths
Maths and puzzle benchmarks reward multi-step thinking because the final answer is easy to verify: OTIS Mock AIME (competition-style problems), FrontierMath T1-3 (original research-level problems), ARC-AGI-2 (abstract puzzles that need a new rule inferred from a few examples) and SimpleBench (trick questions that test common sense).
Coding and software engineering
Coding benchmarks run the code. SWE-bench Verified asks a model to resolve real GitHub issues and passes only if the project's tests pass; Aider Polyglot tests code edits in six languages; Terminal-Bench 2.0 tests command-line work. These scores depend on the agent scaffold around the model as well as the model itself.
Agents and tool use
Agent benchmarks measure whether a model can finish long tasks by calling tools, browsing or operating a computer: MCP Atlas, APEX-Agents, OSWorld 2. They are noisier than question-answer tests because small differences in setup change the outcome.
Human preference
Arena-style benchmarks such as LMArena Text Elo and WebDev Arena do not use an answer key. Real users compare two anonymous answers and vote; the votes become Elo-style ratings. They capture style and helpfulness, not just correctness.
How do you read an LLM benchmark score?
First check the unit. A percentage means the share of tasks solved. An Elo rating has no fixed maximum, and only the gap between two models means anything. A time-based score (for example the length of task a model completes half the time) is a different scale again. Then check the sample size: a board with 20 questions can swing several points by luck. Finally check the date: scores for a model measured last year may not reflect its current version.
Why do different sites report different scores for the same model?
- Different harness. The same model scores differently depending on the prompt, the tools and the agent scaffold.
- Different reasoning effort. Many models run at several effort settings (low, medium, high). The leaderboard keeps the highest-effort row by default.
- Different subset. SWE-bench Verified, Lite and the full test set are different samples, so their numbers are not interchangeable.
- Self-reported versus independent. A lab's own chart and an independent run often disagree. LLMs Tiger prefers the higher-trust source and shows which one it used.
What are benchmark contamination and saturation?
Contamination happens when test questions leak into a model's training data, which inflates scores. Benchmarks such as LiveBench Coding refresh their questions regularly, and SWE-rebench uses newly published tasks, to limit it. Saturation is the opposite problem: once most strong models score near 100%, the benchmark cannot rank them. Both are reasons to look at several benchmarks rather than one.
Which benchmark should you look at for your use case?
| If you need | Look at |
|---|---|
| Code generation and bug fixing | SWE-bench Verified, Aider Polyglot, Terminal-Bench 2.0; see best LLMs for coding |
| Hard reasoning and science | GPQA Diamond, Humanity's Last Exam, ARC-AGI-2; see best LLMs for reasoning |
| Agents and tool use | MCP Atlas, APEX-Agents; see best LLMs for agents |
| A pleasant chat assistant | LMArena Text Elo |
| Low cost at good quality | cheapest strong LLM APIs |
| Weights you can run yourself | best open-source LLMs |
What are the limits of LLM benchmarks?
A benchmark is a proxy. It tells you how a model does on a defined test, not how it will behave on your documents, your prompts or your users. Treat the leaderboard as a shortlist, then test the two or three best candidates on a sample of your own work before you commit.
How LLMs Tiger helps
We collect scores from public, cited sources, keep the source and capture date next to every number, and combine them into one aggregate score so you can compare models quickly. How the aggregate works, what is excluded and how often data refreshes is documented on the methodology page. Browse all benchmark leaderboards or the top 100 models.