Claude Opus 4.7 (max): benchmark scores, price and ranking

Claude Opus 4.7 (max) from Anthropic ranks #33 of 741 LLMs on the LLMs Tiger aggregate benchmark score (53.7), based on 19 benchmark results captured up to 2026-10-05.

Organisation
Anthropic
Type
Frontier model
Open weights
No
Released
2026-04-16 (source: epoch)
Cheapest API price
$5.00 in / $25.00 out per million tokens (via anthropic, observed 2026-10-05T21:47Z)
Aggregate rank
#33 of 741 (score 53.7)

Relative to other models, Claude Opus 4.7 (max) is strongest on SWE-bench Verified (#1 of 73) and weakest on Furniture Assembly (#18 of 25).

Claude Opus 4.7 (max) benchmark results

Claude Opus 4.7 (max) benchmark scores, ranked against all models measured on each benchmark
BenchmarkScoreRankSourceCaptured
SWE-bench Verified83.5%#1 of 73Epoch AI Benchmarking Hub2026-10-05
DTBench94.7%#13 of 149Epoch AI Benchmarking Hub2026-10-05
WeirdML75.5%#12 of 112Epoch AI Benchmarking Hub2026-10-05
LMCA52.2%#12 of 111Epoch AI Benchmarking Hub2026-10-05
MCP Atlas79.1%#2 of 14Scale Labs MCP Atlas leaderboard2026-10-05
SciCode54.5%#20 of 100Epoch AI Benchmarking Hub2026-10-05
GPQA Diamond86.4%#65 of 274Epoch AI Benchmarking Hub2026-10-05
ProofBench54.0%#11 of 43Epoch AI Benchmarking Hub2026-10-05
OTIS Mock AIME86.7%#49 of 188Epoch AI Benchmarking Hub2026-10-05
FrontierMath T1-370.2%#21 of 69Epoch AI Benchmarking Hub2026-10-05
CritPt12.0%#34 of 111Epoch AI Benchmarking Hub2026-10-05
OSWorld 218.2%#3 of 8Epoch AI Benchmarking Hub2026-10-05
Mystery Games28.0%#24 of 57Epoch AI Benchmarking Hub2026-10-05
FrontierMath T431.7%#25 of 52Epoch AI Benchmarking Hub2026-10-05
GDP.pdf21.0%#12 of 22Epoch AI Benchmarking Hub2026-10-05
APEX-Agents49.2%#8 of 14Epoch AI Benchmarking Hub2026-10-05
Chess Puzzles7.0%#72 of 124Epoch AI Benchmarking Hub2026-10-05
EBR-Bench19.0%#13 of 19Epoch AI Benchmarking Hub2026-10-05
Furniture Assembly33.3%#18 of 25Epoch AI Benchmarking Hub2026-10-05

Claude Opus 4.7 (max) is also listed at 7 other reasoning-effort settings; this page shows the highest-effort row, which is the one the leaderboard keeps by default.

Models ranked nearby

Compare it on the live leaderboard → · How the aggregate score works