WeirdML leaderboard: which LLMs score highest

WeirdML (unusual machine-learning tasks) currently has scores for 112 models. The top score is 93.6% by GPT-6 Astra (pro, max); the median model scores 41.8% and the lowest scores 1.7%.

WeirdML leaderboard: top 50 of 112 models (highest-effort row per model)
#ModelOrganisationWeirdML scoreCheapest input $/MSourceCaptured
1GPT-6 Astra (pro, max)OpenAI93.6%$10.00Epoch AI Benchmarking Hub2026-10-05
2GPT-6 Astra (max)OpenAI93.3%$10.00Epoch AI Benchmarking Hub2026-10-05
3Claude Fable 5.1 (max)Anthropic92.9%$10.00Epoch AI Benchmarking Hub2026-10-05
4Claude Fable 5 (max)Anthropic91.9%$10.00Epoch AI Benchmarking Hub2026-10-05
5Claude Opus 5 (max)Anthropic91.8%$5.00Epoch AI Benchmarking Hub2026-10-05
6GPT-5.6 Sol (pro, max)OpenAI89.4%$2.00Epoch AI Benchmarking Hub2026-10-05
7GPT-5.6 Sol (max)OpenAI87.0%$2.00Epoch AI Benchmarking Hub2026-10-05
8GPT-5.5 (xhigh)OpenAI84.9%$5.00Epoch AI Benchmarking Hub2026-10-05
9Kimi K3 (max)Moonshot82.6%$1.29Epoch AI Benchmarking Hub2026-10-05
10GPT-5.3 Codex (xhigh)OpenAI77.9%$1.75Epoch AI Benchmarking Hub2026-10-05
11GPT-5.4 (xhigh)OpenAI77.7%$2.50Epoch AI Benchmarking Hub2026-10-05
12Claude Opus 4.7 (max)Anthropic75.5%$5.00Epoch AI Benchmarking Hub2026-10-05
13GLM-5.3 (max)Z.ai (Zhipu AI)75.4%$0.070Epoch AI Benchmarking Hub2026-10-05
14GPT-5.2 (xhigh)OpenAI72.2%$1.75Epoch AI Benchmarking Hub2026-10-05
15Gemini 3 Pro PreviewGoogle DeepMind69.9%$2.00Epoch AI Benchmarking Hub2026-10-05
16DeepSeek V4 Pro 0813 (max)DeepSeek66.2%$0.660Epoch AI Benchmarking Hub2026-10-05
17DeepSeek V4 Flash 0731 (max)DeepSeek63.0%$0.015Epoch AI Benchmarking Hub2026-10-05
18GPT-5.1 (high)OpenAI60.8%$1.25Epoch AI Benchmarking Hub2026-10-05
19GPT-5 (high)OpenAI60.7%$1.25Epoch AI Benchmarking Hub2026-10-05
20GPT-5 ProOpenAI60.4%$15.00Epoch AI Benchmarking Hub2026-10-05
21Muse Spark 1.2 (xhigh)Meta AI60.3%$1.25Epoch AI Benchmarking Hub2026-10-05
22o3-pro-2025-06-10 (high)OpenAI58.2%$20.00Epoch AI Benchmarking Hub2026-10-05
23GLM-5.1Z.ai (Zhipu AI)57.1%$1.05Epoch AI Benchmarking Hub2026-10-05
24Kimi K2.6Moonshot55.9%$0.650Epoch AI Benchmarking Hub2026-10-05
25GPT-5-codex (high)OpenAI54.5%$1.25Epoch AI Benchmarking Hub2026-10-05
26Kimi K2.7 CodeMoonshot54.1%$0.671Epoch AI Benchmarking Hub2026-10-05
27gemini-2.5-pro_16KGoogle DeepMind54.0%—Epoch AI Benchmarking Hub2026-10-05
28GPT-5 mini (high)OpenAI52.7%$0.250Epoch AI Benchmarking Hub2026-10-05
29o4-mini (high)OpenAI52.6%$1.00Epoch AI Benchmarking Hub2026-10-05
30o3 (high)OpenAI52.4%$2.00Epoch AI Benchmarking Hub2026-10-05
31Gemma 4 31B ITGoogle DeepMind52.3%$0.100Epoch AI Benchmarking Hub2026-10-05
32grok-4-20xAI52.3%$1.25Epoch AI Benchmarking Hub2026-10-05
33DeepSeek v4 Pro (max)DeepSeek48.9%$0.435Epoch AI Benchmarking Hub2026-10-05
34GLM-5Z.ai (Zhipu AI)48.2%$0.600Epoch AI Benchmarking Hub2026-10-05
35gpt-oss-120b (high)OpenAI48.2%$0.030Epoch AI Benchmarking Hub2026-10-05
36o1-previewOpenAI47.6%$15.00Epoch AI Benchmarking Hub2026-10-05
37DeepSeek-V3.2-SpecialeDeepSeek46.7%$0.580Epoch AI Benchmarking Hub2026-10-05
38claude-sonnet-4-20250514_16KAnthropic46.1%—Epoch AI Benchmarking Hub2026-10-05
39o1 (high)OpenAI46.1%$15.00Epoch AI Benchmarking Hub2026-10-05
40claude-opus-4-1-20250805_16KAnthropic45.9%—Epoch AI Benchmarking Hub2026-10-05
41grok-4-0709xAI45.7%—Epoch AI Benchmarking Hub2026-10-05
42DeepSeek v4 Flash (max)DeepSeek45.6%$0.090Epoch AI Benchmarking Hub2026-10-05
43Kimi K2.5Moonshot45.6%$0.450Epoch AI Benchmarking Hub2026-10-05
44claude-haiku-4-5-20251001Anthropic45.4%$1.00Epoch AI Benchmarking Hub2026-10-05
45claude-haiku-4-5-20251001_16KAnthropic44.1%—Epoch AI Benchmarking Hub2026-10-05
46claude-sonnet-4-20250514Anthropic43.9%—Epoch AI Benchmarking Hub2026-10-05
47Claude Opus 4Anthropic43.7%$15.00Epoch AI Benchmarking Hub2026-10-05
48Mistral Medium 3.5Mistral AI43.7%$1.50Epoch AI Benchmarking Hub2026-10-05
49o3-mini (high)OpenAI43.7%$1.10Epoch AI Benchmarking Hub2026-10-05
50nemotron-3-ultraNvidia43.5%$0.500Epoch AI Benchmarking Hub2026-10-05

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to Kimi K3 (max) at $1.29 per million input tokens (score 82.6%).

What WeirdML measures

WeirdML measures unusual machine-learning tasks.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-05; the Captured column gives each row's date.

Related benchmarks