CL-bench leaderboard: which LLMs score highest

CL-bench (learning new rules from context) currently has scores for 16 models. The top score is 27.9% by GPT-5.4 (xhigh); the median model scores 17.7% and the lowest scores 11.4%.

CL-bench leaderboard: top 16 of 16 models (highest-effort row per model)
#ModelOrganisationCL-bench scoreCheapest input $/MSourceCaptured
1GPT-5.4 (xhigh)OpenAI27.9%$2.50Epoch AI Benchmarking Hub2026-10-05
2GPT-5.1 (high)OpenAI23.7%$1.25Epoch AI Benchmarking Hub2026-10-05
3grok-4-20xAI22.2%$1.25Epoch AI Benchmarking Hub2026-10-05
4Qwen 3.6 Plus (2026-04-02)Alibaba20.3%—Epoch AI Benchmarking Hub2026-10-05
5Qwen3.5 PlusAlibaba19.8%—Epoch AI Benchmarking Hub2026-10-05
6Kimi K2.5Moonshot19.3%$0.450Epoch AI Benchmarking Hub2026-10-05
7GLM-5Z.ai (Zhipu AI)18.7%$0.600Epoch AI Benchmarking Hub2026-10-05
8o3 (high)OpenAI17.8%$2.00Epoch AI Benchmarking Hub2026-10-05
9Kimi K2 Thinking (Together)Moonshot17.6%$0.550Epoch AI Benchmarking Hub2026-10-05
10GLM-4.7 (Novita)Z.ai (Zhipu AI)15.9%$0.400Epoch AI Benchmarking Hub2026-10-05
11Gemini 3 Pro PreviewGoogle DeepMind15.8%$2.00Epoch AI Benchmarking Hub2026-10-05
12mimo-v2-proXiaomi Corp15.7%$1.10Epoch AI Benchmarking Hub2026-10-05
13Qwen3-Max-InstructAlibaba14.5%—Epoch AI Benchmarking Hub2026-10-05
14DeepSeek-V3.2 (Thinking; Novita)DeepSeek12.4%$0.259Epoch AI Benchmarking Hub2026-10-05
15Kimi K2 ThinkingMoonshot11.9%$0.550Epoch AI Benchmarking Hub2026-10-05
16MiniMax-M2.5MiniMax11.4%$0.270Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to GLM-4.7 (Novita) at $0.400 per million input tokens (score 15.9%).

What CL-bench measures

CL-bench measures learning new rules from context.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks