CL-bench leaderboard: which LLMs score highest
CL-bench (learning new rules from context) currently has scores for 16 models. The top score is 27.9% by GPT-5.4 (xhigh); the median model scores 17.7% and the lowest scores 11.4%.
| # | Model | Organisation | CL-bench score | Cheapest input $/M | Source | Captured |
|---|---|---|---|---|---|---|
| 1 | GPT-5.4 (xhigh) | OpenAI | 27.9% | $2.50 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 2 | GPT-5.1 (high) | OpenAI | 23.7% | $1.25 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 3 | grok-4-20 | xAI | 22.2% | $1.25 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 4 | Qwen 3.6 Plus (2026-04-02) | Alibaba | 20.3% | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 5 | Qwen3.5 Plus | Alibaba | 19.8% | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 6 | Kimi K2.5 | Moonshot | 19.3% | $0.450 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 7 | GLM-5 | Z.ai (Zhipu AI) | 18.7% | $0.600 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 8 | o3 (high) | OpenAI | 17.8% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 9 | Kimi K2 Thinking (Together) | Moonshot | 17.6% | $0.550 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 10 | GLM-4.7 (Novita) | Z.ai (Zhipu AI) | 15.9% | $0.400 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 11 | Gemini 3 Pro Preview | Google DeepMind | 15.8% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 12 | mimo-v2-pro | Xiaomi Corp | 15.7% | $1.10 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 13 | Qwen3-Max-Instruct | Alibaba | 14.5% | — | Epoch AI Benchmarking Hub | 2026-10-05 |
| 14 | DeepSeek-V3.2 (Thinking; Novita) | DeepSeek | 12.4% | $0.259 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 15 | Kimi K2 Thinking | Moonshot | 11.9% | $0.550 | Epoch AI Benchmarking Hub | 2026-10-05 |
| 16 | MiniMax-M2.5 | MiniMax | 11.4% | $0.270 | Epoch AI Benchmarking Hub | 2026-10-05 |
Compare all benchmarks side by side on the live leaderboard →
Among the ten highest scorers, the cheapest listed API price belongs to GLM-4.7 (Novita) at $0.400 per million input tokens (score 15.9%).
What CL-bench measures
CL-bench measures learning new rules from context.
How to read these scores
Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.
Where the data comes from
Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.