APEX-Agents leaderboard: which LLMs score highest
APEX-Agents (long professional agent tasks) currently has scores for 17 models. The top score is 75.5% by Claude Sonnet 5.5 (max); the median model scores 49.2% and the lowest scores 21.3%.
| # | Model | Organisation | APEX-Agents score | Cheapest input $/M | Source | Captured |
|---|---|---|---|---|---|---|
| 1 | Claude Sonnet 5.5 (max) | Anthropic | 75.5% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-06 |
| 2 | Claude Opus 5.5 (max) | Anthropic | 73.5% | $4.00 | Epoch AI Benchmarking Hub | 2026-10-06 |
| 3 | Claude Opus 5 (max) | Anthropic | 65.8% | $5.00 | Epoch AI Benchmarking Hub | 2026-10-06 |
| 4 | GPT-6.1 Sol (max) | OpenAI | 60.0% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-06 |
| 5 | mimo-v2.6-pro | Xiaomi Corp | 59.5% | $0.435 | Epoch AI Benchmarking Hub | 2026-10-06 |
| 6 | GPT-5.6 Terra (max) | OpenAI | 58.2% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-06 |
| 7 | GPT-6 Sol (max) | OpenAI | 54.3% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-06 |
| 8 | GPT-5.6 Sol (pro, max) | OpenAI | 51.4% | $2.00 | Epoch AI Benchmarking Hub | 2026-10-06 |
| 9 | Claude Opus 4.7 (max) | Anthropic | 49.2% | $5.00 | Epoch AI Benchmarking Hub | 2026-10-06 |
| 10 | Claude Opus 4.6 (max) | Anthropic | 46.3% | $5.00 | Epoch AI Benchmarking Hub | 2026-10-06 |
| 11 | GPT-6 Luna (max) | OpenAI | 44.3% | $0.100 | Epoch AI Benchmarking Hub | 2026-10-06 |
| 12 | GLM-5.1 | Z.ai (Zhipu AI) | 40.9% | $1.05 | Epoch AI Benchmarking Hub | 2026-10-06 |
| 13 | MiniMax-M3 | MiniMax | 37.7% | $0.230 | Epoch AI Benchmarking Hub | 2026-10-06 |
| 14 | Kimi K2.7 Code | Moonshot | 37.6% | $0.671 | Epoch AI Benchmarking Hub | 2026-10-06 |
| 15 | Qwen3.5 397B-A17B | Alibaba | 24.9% | $0.164 | Epoch AI Benchmarking Hub | 2026-10-06 |
| 16 | nemotron-3-ultra | Nvidia | 22.7% | $0.500 | Epoch AI Benchmarking Hub | 2026-10-06 |
| 17 | DeepSeek-V3.2 (Thinking; Novita) | DeepSeek | 21.3% | $0.260 | Epoch AI Benchmarking Hub | 2026-10-06 |
Compare all benchmarks side by side on the live leaderboard →
Among the ten highest scorers, the cheapest listed API price belongs to mimo-v2.6-pro at $0.435 per million input tokens (score 59.5%).
What APEX-Agents measures
APEX-Agents measures long professional agent tasks.
How to read these scores
Scores are percentages (higher is better): the share of questions or tasks the model completed correctly. Where a model is listed at several reasoning efforts, only its highest-effort row is shown. See the methodology for how rows are chosen.
Where the data comes from
Epoch AI Benchmarking Hub (CC BY 4.0). Scores were last captured 2026-10-06; the Captured column gives each row's date.