GPT-5.5 (xhigh): benchmark scores, price and ranking

GPT-5.5 (xhigh) from OpenAI ranks #10 of 741 LLMs on the LLMs Tiger aggregate benchmark score (58.3), based on 28 benchmark results captured up to 2026-10-05.

Organisation
OpenAI
Type
Frontier model
Open weights
No
Released
2026-04-23 (source: epoch)
Cheapest API price
$5.00 in / $30.00 out per million tokens (via aihubmix, observed 2026-10-05T21:47Z)
Aggregate rank
#10 of 741 (score 58.3)

Relative to other models, GPT-5.5 (xhigh) is strongest on GSO-Bench (#1 of 19) and weakest on DeepSWE (#10 of 19).

GPT-5.5 (xhigh) benchmark results

GPT-5.5 (xhigh) benchmark scores, ranked against all models measured on each benchmark
BenchmarkScoreRankSourceCaptured
GSO-Bench40.2%#1 of 19Epoch AI Benchmarking Hub2026-10-05
DTBench96.0%#8 of 149Epoch AI Benchmarking Hub2026-10-05
LiveBench Data Analysis81.6%#1 of 17LiveBench 2026_06_25 official CSV2026-10-05
LiveBench Language87.4%#1 of 17LiveBench 2026_06_25 official CSV2026-10-05
LiveBench Mathematics95.9%#1 of 17LiveBench 2026_06_25 official CSV2026-10-05
ALE-Bench1,943#5 of 78Epoch AI Benchmarking Hub2026-10-05
WeirdML84.9%#8 of 112Epoch AI Benchmarking Hub2026-10-05
LMCA54.3%#9 of 111Epoch AI Benchmarking Hub2026-10-05
ARC-AGI-285.0%#10 of 89ARC Prize official leaderboard2026-10-05
CritPt27.1%#13 of 111Epoch AI Benchmarking Hub2026-10-05
LiveBench Coding82.1%#2 of 17LiveBench 2026_06_25 official CSV2026-10-05
LiveBench Reasoning89.7%#2 of 17LiveBench 2026_06_25 official CSV2026-10-05
SimpleQA Verified63.0%#8 of 66Epoch AI Benchmarking Hub2026-10-05
Mystery Games56.0%#8 of 57Epoch AI Benchmarking Hub2026-10-05
SciCode56.1%#17 of 100Epoch AI Benchmarking Hub2026-10-05
FrontierMath T1-385.3%#12 of 69Epoch AI Benchmarking Hub2026-10-05
LiveBench Agentic Coding52.1%#3 of 17LiveBench 2026_06_25 official CSV2026-10-05
FrontierMath T472.5%#13 of 52Epoch AI Benchmarking Hub2026-10-05
WebDev Arena1,510#17 of 63Epoch AI Benchmarking Hub2026-10-05
GDP.pdf26.0%#6 of 22Epoch AI Benchmarking Hub2026-10-05
LiveBench IF70.7%#5 of 17LiveBench 2026_06_25 official CSV2026-10-05
ProofBench50.0%#13 of 43Epoch AI Benchmarking Hub2026-10-05
Furniture Assembly44.2%#10 of 25Epoch AI Benchmarking Hub2026-10-05
MCP Atlas75.3%#6 of 14Scale Labs MCP Atlas leaderboard2026-10-05
EBR-Bench34.3%#9 of 19Epoch AI Benchmarking Hub2026-10-05
OSWorld 213.0%#4 of 8Epoch AI Benchmarking Hub2026-10-05
DeepSWE67.0%#10 of 19Epoch AI Benchmarking Hub2026-10-05
PostTrainBench27.2%#3 of 4Epoch AI Benchmarking Hub2026-10-05

GPT-5.5 (xhigh) is also listed at 6 other reasoning-effort settings; this page shows the highest-effort row, which is the one the leaderboard keeps by default.

Models ranked nearby

Compare it on the live leaderboard → · How the aggregate score works