Claude Opus 5.5 (max): benchmark scores, price and ranking

Claude Opus 5.5 (max) from Anthropic ranks #1 of 741 LLMs on the LLMs Tiger aggregate benchmark score (64.2), based on 21 benchmark results captured up to 2026-10-05.

Organisation
Anthropic
Type
Frontier model
Open weights
No
Released
2026-09-22 (source: curated)
Cheapest API price
$4.00 in / $20.00 out per million tokens (via anthropic, observed 2026-10-05T21:47Z)
Aggregate rank
#1 of 741 (score 64.2)

Relative to other models, Claude Opus 5.5 (max) is strongest on DTBench (#1 of 149) and weakest on APEX-Agents (#2 of 14).

Claude Opus 5.5 (max) benchmark results

Claude Opus 5.5 (max) benchmark scores, ranked against all models measured on each benchmark
BenchmarkScoreRankSourceCaptured
DTBench98.9%#1 of 149Epoch AI Benchmarking Hub2026-10-05
LMCA68.2%#1 of 111Epoch AI Benchmarking Hub2026-10-05
SciCode66.9%#1 of 100Epoch AI Benchmarking Hub2026-10-05
OTIS Mock AIME100.0%#2 of 188Epoch AI Benchmarking Hub2026-10-05
WebDev Arena1,820#1 of 63Epoch AI Benchmarking Hub2026-10-05
CritPt31.7%#2 of 111Epoch AI Benchmarking Hub2026-10-05
LMArena Code Elo1,815#1 of 29LMArena2026-10-05
Furniture Assembly83.3%#1 of 25Epoch AI Benchmarking Hub2026-10-05
FrontierMath T1-391.2%#3 of 69Epoch AI Benchmarking Hub2026-10-05
ARC-AGI-291.7%#4 of 89ARC Prize official leaderboard2026-10-05
SimpleQA Verified72.2%#3 of 66Epoch AI Benchmarking Hub2026-10-05
ProofBench100.0%#2 of 43Epoch AI Benchmarking Hub2026-10-05
Mystery Games71.0%#3 of 57Epoch AI Benchmarking Hub2026-10-05
FrontierMath T495.0%#3 of 52Epoch AI Benchmarking Hub2026-10-05
CursorBench57.8%#1 of 13Epoch AI Benchmarking Hub2026-10-05
GDP.pdf30.6%#2 of 22Epoch AI Benchmarking Hub2026-10-05
EBR-Bench71.4%#2 of 19Epoch AI Benchmarking Hub2026-10-05
GPQA Diamond90.6%#32 of 274Epoch AI Benchmarking Hub2026-10-05
FrontierSWE62.3%#2 of 16Epoch AI Benchmarking Hub2026-10-05
APEX-Agents73.5%#2 of 14Epoch AI Benchmarking Hub2026-10-05
MirrorCode77.4%#1 of 1Epoch AI Benchmarking Hub2026-10-05

Claude Opus 5.5 (max) is also listed at 6 other reasoning-effort settings; this page shows the highest-effort row, which is the one the leaderboard keeps by default.

Models ranked nearby

Compare it on the live leaderboard → · How the aggregate score works