GPT-5.4 (xhigh): benchmark scores, price and ranking

GPT-5.4 (xhigh) from OpenAI ranks #13 of 741 LLMs on the LLMs Tiger aggregate benchmark score (57.1), based on 31 benchmark results captured up to 2026-10-05.

Organisation
OpenAI
Type
Frontier model
Open weights
No
Released
2026-03-05 (source: epoch)
Cheapest API price
$2.50 in / $15.00 out per million tokens (via aihubmix, observed 2026-10-05T21:47Z)
Aggregate rank
#13 of 741 (score 57.1)

Relative to other models, GPT-5.4 (xhigh) is strongest on GPQA Diamond (#13 of 274) and weakest on DeepSWE (#18 of 19).

GPT-5.4 (xhigh) benchmark results

GPT-5.4 (xhigh) benchmark scores, ranked against all models measured on each benchmark
BenchmarkScoreRankSourceCaptured
GPQA Diamond93.3%#13 of 274Epoch AI Benchmarking Hub2026-10-05
METR time horizon342 min#2 of 34Epoch AI Benchmarking Hub2026-10-05
CL-bench27.9%#1 of 16Epoch AI Benchmarking Hub2026-10-05
DTBench94.4%#14 of 149Epoch AI Benchmarking Hub2026-10-05
Chess Puzzles44.0%#12 of 124Epoch AI Benchmarking Hub2026-10-05
WeirdML77.7%#11 of 112Epoch AI Benchmarking Hub2026-10-05
Humanity's Last Exam36.2%#3 of 29Epoch AI Benchmarking Hub2026-10-05
GSO-Bench31.4%#2 of 19Epoch AI Benchmarking Hub2026-10-05
CL-bench Life21.7%#1 of 9Epoch AI Benchmarking Hub2026-10-05
EnigmaEval16.0%#3 of 26Epoch AI Benchmarking Hub2026-10-05
LMCA52.0%#13 of 111Epoch AI Benchmarking Hub2026-10-05
LiveBench Agentic Coding53.8%#2 of 17LiveBench 2026_06_25 official CSV2026-10-05
LiveBench Data Analysis79.3%#2 of 17LiveBench 2026_06_25 official CSV2026-10-05
SciCode56.6%#13 of 100Epoch AI Benchmarking Hub2026-10-05
OTIS Mock AIME95.3%#27 of 188Epoch AI Benchmarking Hub2026-10-05
CritPt23.4%#16 of 111Epoch AI Benchmarking Hub2026-10-05
LiveBench Mathematics94.1%#3 of 17LiveBench 2026_06_25 official CSV2026-10-05
LiveBench Reasoning88.1%#3 of 17LiveBench 2026_06_25 official CSV2026-10-05
ARC-AGI-274.0%#18 of 89ARC Prize official leaderboard2026-10-05
FrontierMath T1-378.6%#16 of 69Epoch AI Benchmarking Hub2026-10-05
ProofBench56.0%#10 of 43Epoch AI Benchmarking Hub2026-10-05
LiveBench Language82.6%#4 of 17LiveBench 2026_06_25 official CSV2026-10-05
Mystery Games37.0%#15 of 57Epoch AI Benchmarking Hub2026-10-05
FrontierMath T449.0%#18 of 52Epoch AI Benchmarking Hub2026-10-05
LiveBench IF70.2%#6 of 17LiveBench 2026_06_25 official CSV2026-10-05
SimpleQA Verified45.1%#29 of 66Epoch AI Benchmarking Hub2026-10-05
LiveBench Coding77.5%#8 of 17LiveBench 2026_06_25 official CSV2026-10-05
MCP Atlas70.6%#7 of 14Scale Labs MCP Atlas leaderboard2026-10-05
EBR-Bench25.4%#11 of 19Epoch AI Benchmarking Hub2026-10-05
Furniture Assembly37.5%#15 of 25Epoch AI Benchmarking Hub2026-10-05
DeepSWE51.8%#18 of 19Epoch AI Benchmarking Hub2026-10-05

GPT-5.4 (xhigh) is also listed at 6 other reasoning-effort settings; this page shows the highest-effort row, which is the one the leaderboard keeps by default.

Models ranked nearby

  • #10 GPT-5.5 (xhigh) (58.3)
  • #11 GPT-5.6 Terra (max) (57.7)
  • #12 GPT-5.5 Pro (xhigh) (57.5)
  • #14 claude-opus-4-8-xhigh-effort (56.6)
  • #15 GPT-5.4 Pro (xhigh) (56.1)
  • #16 GPT-5.6 Sol (pro, max) (56.1)

Compare it on the live leaderboard → · How the aggregate score works