Best LLMs for AI Agents and Tool Use (2026)

Claude Opus 5.5 (max) leads this ranking with a score of 57.2 (50 is the neutral midpoint). This page ranks models on agent benchmarks: tool use over MCP servers, long professional tasks, computer use, game playing, research reports and paid freelance-style work. Data last captured 2026-10-05.

Best LLMs for Agents and Tool Use: top 21
#ModelOrganisationScoreOverall rankCheapest input $/M
1Claude Opus 5.5 (max)Anthropic57.2#1$4.00
2Claude Opus 5 (max)Anthropic55.6#6$5.00
3GPT-5.6 Sol (max)OpenAI55.2#7$2.00
4Gemini 3 Pro PreviewGoogle DeepMind54.6#30$2.00
5GPT-5.5 (xhigh)OpenAI53.1#10$5.00
6GPT-6 Sol (max)OpenAI53.0#8$2.00
7GPT-5.6 Terra (max)OpenAI52.8#11$2.00
8Claude Opus 4.7 (max)Anthropic52.1#33$5.00
9Claude Opus 4.6 (max)Anthropic51.5#67$5.00
10GPT-5.6 Luna (max)OpenAI50.7#19$0.200
11GPT-6 Luna (max)OpenAI50.5#34$0.100
12Grok 4.6 (xhigh)xAI50.4#36$1.25
13grok-4-0709xAI50.0#40—
14Claude Sonnet 4.6 (max)Anthropic49.8#212$3.00
15Claude 3.5 Sonnet (Oct 2024)Anthropic49.3#596$3.00
16Kimi K2.6Moonshot48.4#351$0.650
17MiniMax-M3MiniMax47.8#570$0.230
18gemini-2.5-flashGoogle DeepMind47.6#576$0.300
19Llama-3.1-70B-InstructMeta AI46.5#699$0.400
20qwen2.5-72b-instructAlibaba45.6#709$0.120
21gemini-1.5-pro-002Google DeepMind45.4#661—

How this ranking is built

This ranking uses only agent benchmarks: tool use over MCP servers, long professional tasks, computer use, game playing, research reports and paid freelance-style work. For each model we take its z-score on every one of these benchmarks it has been measured on and shrink the average toward 50 when it has few results; models need at least three of them to appear. Benchmarks used: BALROG, DTBench, Terminal-Bench 2.0, Terminal-Bench 4.0 and GDP.pdf.

More rankings