grok-4-0709: benchmark scores, price and ranking

grok-4-0709 from xAI ranks #40 of 741 LLMs on the LLMs Tiger aggregate benchmark score (53.2), based on 11 benchmark results captured up to 2026-10-05.

Organisation
xAI
Type
Frontier model
Open weights
No
Released
2025-07-09 (source: epoch)
Aggregate rank
#40 of 741 (score 53.2)

Relative to other models, grok-4-0709 is strongest on Fiction.liveBench (#1 of 47) and weakest on Terminal-Bench 2.0 (#21 of 28).

grok-4-0709 benchmark results

grok-4-0709 benchmark scores, ranked against all models measured on each benchmark
BenchmarkScoreRankSourceCaptured
Fiction.liveBench94.4%#1 of 47Epoch AI Benchmarking Hub2026-10-05
Cybench43.0%#1 of 16Epoch AI Benchmarking Hub2026-10-05
SimpleBench60.5%#10 of 64Epoch AI Benchmarking Hub2026-10-05
BALROG43.6%#7 of 34Epoch AI Benchmarking Hub2026-10-05
METR time horizon110 min#7 of 34Epoch AI Benchmarking Hub2026-10-05
GPQA Diamond87.0%#59 of 274Epoch AI Benchmarking Hub2026-10-05
Chess Puzzles28.0%#30 of 124Epoch AI Benchmarking Hub2026-10-05
OTIS Mock AIME84.0%#58 of 188Epoch AI Benchmarking Hub2026-10-05
WeirdML45.7%#41 of 112Epoch AI Benchmarking Hub2026-10-05
DeepResearch Bench47.3%#3 of 7Epoch AI Benchmarking Hub2026-10-05
Terminal-Bench 2.027.2%#21 of 28Epoch AI Benchmarking Hub2026-10-05

Models ranked nearby

  • #37 Gemini 3.8 Flash (high) (53.5)
  • #38 gemini-2.5-pro-preview-06-05 (53.5)
  • #39 Qwen3.8 Max (xhigh) (53.2)
  • #41 Grok 4 (53.2)
  • #42 NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 (53.1)
  • #43 NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 (53.1)

Compare it on the live leaderboard → · How the aggregate score works