Claude Sonnet 5.5 (max): benchmark scores, price and ranking

Claude Sonnet 5.5 (max) from Anthropic ranks #4 of 741 LLMs on the LLMs Tiger aggregate benchmark score (60.8), based on 13 benchmark results captured up to 2026-10-05.

Organisation
Anthropic
Type
Frontier model
Open weights
No
Released
2026-09-28 (source: curated)
Cheapest API price
$2.00 in / $10.00 out per million tokens (via anthropic, observed 2026-10-05T21:47Z)
Aggregate rank
#4 of 741 (score 60.8)

Relative to other models, Claude Sonnet 5.5 (max) is strongest on GPQA Diamond (#2 of 274) and weakest on SimpleQA Verified (#26 of 66).

Claude Sonnet 5.5 (max) benchmark results

Claude Sonnet 5.5 (max) benchmark scores, ranked against all models measured on each benchmark
BenchmarkScoreRankSourceCaptured
GPQA Diamond95.6%#2 of 274Epoch AI Benchmarking Hub2026-10-05
OTIS Mock AIME100.0%#3 of 188Epoch AI Benchmarking Hub2026-10-05
SciCode61.0%#4 of 100Epoch AI Benchmarking Hub2026-10-05
CritPt31.4%#5 of 111Epoch AI Benchmarking Hub2026-10-05
ProofBench100.0%#3 of 43Epoch AI Benchmarking Hub2026-10-05
Mystery Games65.0%#4 of 57Epoch AI Benchmarking Hub2026-10-05
APEX-Agents75.5%#1 of 14Epoch AI Benchmarking Hub2026-10-05
FrontierMath T1-388.8%#7 of 69Epoch AI Benchmarking Hub2026-10-05
FrontierMath T480.5%#8 of 52Epoch AI Benchmarking Hub2026-10-05
CursorBench55.5%#2 of 13Epoch AI Benchmarking Hub2026-10-05
Furniture Assembly75.0%#4 of 25Epoch AI Benchmarking Hub2026-10-05
FrontierSWE61.9%#3 of 16Epoch AI Benchmarking Hub2026-10-05
SimpleQA Verified46.5%#26 of 66Epoch AI Benchmarking Hub2026-10-05

Claude Sonnet 5.5 (max) is also listed at 5 other reasoning-effort settings; this page shows the highest-effort row, which is the one the leaderboard keeps by default.

Models ranked nearby

Compare it on the live leaderboard → · How the aggregate score works