APEX-Agents leaderboard: which LLMs score highest

APEX-Agents (long professional agent tasks) currently has scores for 17 models. The top score is 75.5% by Claude Sonnet 5.5 (max); the median model scores 49.2% and the lowest scores 21.3%.

APEX-Agents leaderboard: top 17 of 17 models (highest-effort row per model)
#ModelOrganisationAPEX-Agents scoreCheapest input $/MSourceCaptured
1Claude Sonnet 5.5 (max)Anthropic75.5%$2.00Epoch AI Benchmarking Hub2026-10-06
2Claude Opus 5.5 (max)Anthropic73.5%$4.00Epoch AI Benchmarking Hub2026-10-06
3Claude Opus 5 (max)Anthropic65.8%$5.00Epoch AI Benchmarking Hub2026-10-06
4GPT-6.1 Sol (max)OpenAI60.0%$2.00Epoch AI Benchmarking Hub2026-10-06
5mimo-v2.6-proXiaomi Corp59.5%$0.435Epoch AI Benchmarking Hub2026-10-06
6GPT-5.6 Terra (max)OpenAI58.2%$2.00Epoch AI Benchmarking Hub2026-10-06
7GPT-6 Sol (max)OpenAI54.3%$2.00Epoch AI Benchmarking Hub2026-10-06
8GPT-5.6 Sol (pro, max)OpenAI51.4%$2.00Epoch AI Benchmarking Hub2026-10-06
9Claude Opus 4.7 (max)Anthropic49.2%$5.00Epoch AI Benchmarking Hub2026-10-06
10Claude Opus 4.6 (max)Anthropic46.3%$5.00Epoch AI Benchmarking Hub2026-10-06
11GPT-6 Luna (max)OpenAI44.3%$0.100Epoch AI Benchmarking Hub2026-10-06
12GLM-5.1Z.ai (Zhipu AI)40.9%$1.05Epoch AI Benchmarking Hub2026-10-06
13MiniMax-M3MiniMax37.7%$0.230Epoch AI Benchmarking Hub2026-10-06
14Kimi K2.7 CodeMoonshot37.6%$0.671Epoch AI Benchmarking Hub2026-10-06
15Qwen3.5 397B-A17BAlibaba24.9%$0.164Epoch AI Benchmarking Hub2026-10-06
16nemotron-3-ultraNvidia22.7%$0.500Epoch AI Benchmarking Hub2026-10-06
17DeepSeek-V3.2 (Thinking; Novita)DeepSeek21.3%$0.260Epoch AI Benchmarking Hub2026-10-06

Compare all benchmarks side by side on the live leaderboard →

Among the ten highest scorers, the cheapest listed API price belongs to mimo-v2.6-pro at $0.435 per million input tokens (score 59.5%).

What APEX-Agents measures

APEX-Agents measures long professional agent tasks.

How to read these scores

Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.

Where the data comes from

Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-06; the Captured column gives each row's date.

Related benchmarks