Grok 4.7 (xhigh): benchmark scores, price and ranking

Grok 4.7 (xhigh) from xAI ranks #45 of 751 LLMs on the LLMs Tiger aggregate benchmark score (53.1), based on 19 benchmark results captured up to 2026-10-07.

Organisation
xAI
Type
Frontier model
Open weights
No
Released
2026-09-21 (source: curated)
Cheapest API price
$2.00 in / $6.00 out per million tokens (via azure_ai, observed 2026-10-08T00:16Z)
Aggregate rank
#45 of 751 (score 53.1)

Relative to other models, Grok 4.7 (xhigh) is strongest on GPQA Diamond (#17 of 274) and weakest on Furniture Assembly (#24 of 25).

Grok 4.7 (xhigh) benchmark results

Grok 4.7 (xhigh) benchmark scores, ranked against all models measured on each benchmark
BenchmarkScoreRankSourceCaptured
GPQA Diamond92.7%#17 of 274Epoch AI Benchmarking Hub2026-10-07
DTBench96.0%#11 of 149Epoch AI Benchmarking Hub2026-10-07
OTIS Mock AIME98.1%#18 of 188Epoch AI Benchmarking Hub2026-10-07
SciCode57.4%#12 of 106Epoch AI Benchmarking Hub2026-10-07
LMCA49.4%#15 of 111Epoch AI Benchmarking Hub2026-10-07
Chess Puzzles38.0%#18 of 124Epoch AI Benchmarking Hub2026-10-07
WebDev Arena1,638#10 of 67Epoch AI Benchmarking Hub2026-10-07
SimpleQA Verified56.0%#12 of 66Epoch AI Benchmarking Hub2026-10-07
CritPt17.7%#26 of 117Epoch AI Benchmarking Hub2026-10-07
ARC-AGI-261.4%#30 of 92ARC Prize official leaderboard2026-10-07
LMArena Code Elo1,638#10 of 30LMArena2026-10-07
CursorBench46.3%#5 of 15Epoch AI Benchmarking Hub2026-10-07
WeirdML v38.6%#1 of 3Epoch AI Benchmarking Hub2026-10-07
Mystery Games29.0%#22 of 57Epoch AI Benchmarking Hub2026-10-07
GDP.pdf22.8%#13 of 25Epoch AI Benchmarking Hub2026-10-07
FrontierSWE29.5%#10 of 19Epoch AI Benchmarking Hub2026-10-07
FrontierMath T1-353.0%#38 of 69Epoch AI Benchmarking Hub2026-10-07
FrontierMath T417.1%#39 of 52Epoch AI Benchmarking Hub2026-10-07
Furniture Assembly20.8%#24 of 25Epoch AI Benchmarking Hub2026-10-07

Grok 4.7 (xhigh) is also listed at 5 other reasoning-effort settings; this page shows the highest-effort row, which is the one the leaderboard keeps by default.

Models ranked nearby

  • #42 grok-4-0709 (53.2)
  • #43 Grok 4 (53.2)
  • #44 NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 (53.1)
  • #46 NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 (53.1)
  • #47 o3 + gpt-4.1 (53.0)
  • #48 Claude Code + Sonnet 5.5 (53.0)

Compare it on the live leaderboard → · How the aggregate score works