grok-4-0709: benchmark scores, price and ranking
grok-4-0709 from xAI ranks #40 of 741 LLMs on the LLMs Tiger aggregate benchmark score (53.2), based on 11 benchmark results captured up to 2026-10-05.
- Organisation
- xAI
- Type
- Frontier model
- Open weights
- No
- Released
- 2025-07-09 (source: epoch)
- Aggregate rank
- #40 of 741 (score 53.2)
Relative to other models, grok-4-0709 is strongest on Fiction.liveBench (#1 of 47) and weakest on Terminal-Bench 2.0 (#21 of 28).
grok-4-0709 benchmark results
| Benchmark | Score | Rank | Source | Captured |
|---|---|---|---|---|
| Fiction.liveBench | 94.4% | #1 of 47 | Epoch AI Benchmarking Hub | 2026-10-05 |
| Cybench | 43.0% | #1 of 16 | Epoch AI Benchmarking Hub | 2026-10-05 |
| SimpleBench | 60.5% | #10 of 64 | Epoch AI Benchmarking Hub | 2026-10-05 |
| BALROG | 43.6% | #7 of 34 | Epoch AI Benchmarking Hub | 2026-10-05 |
| METR time horizon | 110 min | #7 of 34 | Epoch AI Benchmarking Hub | 2026-10-05 |
| GPQA Diamond | 87.0% | #59 of 274 | Epoch AI Benchmarking Hub | 2026-10-05 |
| Chess Puzzles | 28.0% | #30 of 124 | Epoch AI Benchmarking Hub | 2026-10-05 |
| OTIS Mock AIME | 84.0% | #58 of 188 | Epoch AI Benchmarking Hub | 2026-10-05 |
| WeirdML | 45.7% | #41 of 112 | Epoch AI Benchmarking Hub | 2026-10-05 |
| DeepResearch Bench | 47.3% | #3 of 7 | Epoch AI Benchmarking Hub | 2026-10-05 |
| Terminal-Bench 2.0 | 27.2% | #21 of 28 | Epoch AI Benchmarking Hub | 2026-10-05 |
Models ranked nearby
- #37 Gemini 3.8 Flash (high) (53.5)
- #38 gemini-2.5-pro-preview-06-05 (53.5)
- #39 Qwen3.8 Max (xhigh) (53.2)
- #41 Grok 4 (53.2)
- #42 NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 (53.1)
- #43 NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 (53.1)
Compare it on the live leaderboard → · How the aggregate score works