LiveBench IF leaderboard: which LLMs score highest
LiveBench IF (LiveBench release 2026_06_25) currently has scores for 17 models. The top score is 79.1% by gemini-3.1-pro-preview-high; the median model scores 65.2% and the lowest scores 53.2%.
| # | Model | Organisation | LiveBench IF score | Cheapest input $/M | Source | Captured |
|---|---|---|---|---|---|---|
| 1 | gemini-3.1-pro-preview-high | 79.1% | $2.00 | LiveBench 2026_06_25 official CSV | 2026-10-05 | |
| 2 | gemini-3.5-flash-high | 75.6% | $1.50 | LiveBench 2026_06_25 official CSV | 2026-10-05 | |
| 3 | Qwen3.7 Max | Alibaba | 74.0% | $1.25 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 4 | claude-opus-4-8-xhigh-effort | Anthropic | 72.4% | $5.00 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 5 | GPT-5.5 (xhigh) | OpenAI | 70.7% | $5.00 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 6 | GPT-5.4 (xhigh) | OpenAI | 70.2% | $2.50 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 7 | GPT-5.4 nano (xhigh) | OpenAI | 67.2% | $0.200 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 8 | GPT-5.2 Codex | OpenAI | 66.4% | $1.75 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 9 | grok-build-0.1 | xAI | 65.2% | $1.00 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 10 | kimi-k2.6-thinking | Moonshot AI | 64.4% | $0.650 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 11 | claude-opus-4-5-20251101-thinking-64k-high-effort | Anthropic | 62.5% | $5.00 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 12 | gpt-5.2-2025-12-11-high | OpenAI | 61.8% | $1.75 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 13 | GPT-5.4 mini (xhigh) | OpenAI | 59.8% | $0.750 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 14 | qwen3.6-plus | Alibaba | 58.3% | $0.325 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 15 | MiniMax-M3 | MiniMax | 57.5% | $0.230 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 16 | Kimi K2.7 Code | Moonshot | 56.3% | $0.671 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
| 17 | Qwen3.6 27B | Alibaba | 53.2% | $0.150 | LiveBench 2026_06_25 official CSV | 2026-10-05 |
Compare all benchmarks side by side on the live leaderboard →
Among the ten highest scorers, the cheapest listed API price belongs to GPT-5.4 nano (xhigh) at $0.200 per million input tokens (score 67.2%).
What LiveBench IF measures
LiveBench IF measures LiveBench release 2026_06_25.
How to read these scores
Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.
Where the data comes from
LiveBench 2026_06_25 official CSV. Scores were last captured 2026-10-05; the Captured column gives each row's date.