Best LLMs for AI Agents and Tool Use (2026)
Claude Opus 5.5 (max) leads this ranking with a score of 57.2 (50 is the neutral midpoint). This page ranks models on agent benchmarks: tool use over MCP servers, long professional tasks, computer use, game playing, research reports and paid freelance-style work. Data last captured 2026-10-05.
| # | Model | Organisation | Score | Overall rank | Cheapest input $/M |
|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 (max) | Anthropic | 57.2 | #1 | $4.00 |
| 2 | Claude Opus 5 (max) | Anthropic | 55.6 | #6 | $5.00 |
| 3 | GPT-5.6 Sol (max) | OpenAI | 55.2 | #7 | $2.00 |
| 4 | Gemini 3 Pro Preview | Google DeepMind | 54.6 | #30 | $2.00 |
| 5 | GPT-5.5 (xhigh) | OpenAI | 53.1 | #10 | $5.00 |
| 6 | GPT-6 Sol (max) | OpenAI | 53.0 | #8 | $2.00 |
| 7 | GPT-5.6 Terra (max) | OpenAI | 52.8 | #11 | $2.00 |
| 8 | Claude Opus 4.7 (max) | Anthropic | 52.1 | #33 | $5.00 |
| 9 | Claude Opus 4.6 (max) | Anthropic | 51.5 | #67 | $5.00 |
| 10 | GPT-5.6 Luna (max) | OpenAI | 50.7 | #19 | $0.200 |
| 11 | GPT-6 Luna (max) | OpenAI | 50.5 | #34 | $0.100 |
| 12 | Grok 4.6 (xhigh) | xAI | 50.4 | #36 | $1.25 |
| 13 | grok-4-0709 | xAI | 50.0 | #40 | — |
| 14 | Claude Sonnet 4.6 (max) | Anthropic | 49.8 | #212 | $3.00 |
| 15 | Claude 3.5 Sonnet (Oct 2024) | Anthropic | 49.3 | #596 | $3.00 |
| 16 | Kimi K2.6 | Moonshot | 48.4 | #351 | $0.650 |
| 17 | MiniMax-M3 | MiniMax | 47.8 | #570 | $0.230 |
| 18 | gemini-2.5-flash | Google DeepMind | 47.6 | #576 | $0.300 |
| 19 | Llama-3.1-70B-Instruct | Meta AI | 46.5 | #699 | $0.400 |
| 20 | qwen2.5-72b-instruct | Alibaba | 45.6 | #709 | $0.120 |
| 21 | gemini-1.5-pro-002 | Google DeepMind | 45.4 | #661 | — |
How this ranking is built
This ranking uses only agent benchmarks: tool use over MCP servers, long professional tasks, computer use, game playing, research reports and paid freelance-style work. For each model we take its z-score on every one of these benchmarks it has been measured on and shrink the average toward 50 when it has few results; models need at least three of them to appear. Benchmarks used: BALROG, DTBench, Terminal-Bench 2.0, Terminal-Bench 4.0 and GDP.pdf.