Methodology: how LLMs Tiger collects and ranks LLM benchmarks

Every score on LLMs Tiger comes from a public source, keeps its source and capture date, and is ranked by one documented formula. We do not run our own tests, and affiliate links never influence a ranking.

Where the scores come from

75 benchmarks are collected from the sources below. Each score on the site links to the page it came from.

One score per model and benchmark

If several sources report the same benchmark for the same model, we keep the one with the highest trust level (official or independent leaderboards rank above self-reported figures). The source used is always shown.

The aggregate score

For each benchmark, every model's score is converted to a z-score: how many standard deviations it sits above or below the typical model on that benchmark. A model's aggregate is 50 plus ten times the weighted average of its z-scores, shrunk toward the average (50) when it has few results, so one lucky benchmark cannot carry a model to the top. A benchmark counts fully once at least 16 models have a score and is ignored below 5. Each z-score is capped at four standard deviations. The result is relative: 50 is the neutral midpoint, not a percentage and not exactly the average (thinly measured models are pulled toward 50, so the typical model sits a little below it).

Some benchmarks are shown as columns but left out of the aggregate because they are narrow or outlier-prone, or duplicate another input: Chess Puzzles, Epoch ECI, LMArena Document Elo, LMArena Search Elo, LMArena Vision Elo, METR time horizon, Mystery Games and Vending-Bench 2.

Best effort only

Many models appear several times at different reasoning-effort settings. By default the leaderboard keeps only the highest-effort row of each model; the lower rows can be switched back on. The pages on this site follow the same rule.

Prices

The price shown is the cheapest listed input price per million tokens we found across providers for that model, with the matching output price. Prices are re-checked every hour and each carries its observed time. They are a convenience, not a quote: always confirm with the provider.

Release dates

Release dates come from a curated list first, then the model name, then the source dataset, then the provider's listing date. If none is known the site shows a dash rather than a guess. Dates are coloured from green (newest) to red (oldest).

Refresh schedule

Scores refresh every 12 hours; prices every hour. The newest score capture on the site is dated 2026-10-05.

Affiliate links

Some provider links in the separate "Try these models" block may be affiliate links; they are labelled as such. The provider table is never used in a ranking query, so money cannot change a rank.

Limits

Benchmarks are proxies for real work. Self-reported and independent runs differ, scaffolds matter, and new models appear faster than some boards update. Read the guide to LLM benchmarks before you rely on a single number.