DeepSeek V4 Flash 0731 (max): benchmark scores, price and ranking
DeepSeek V4 Flash 0731 (max) from DeepSeek ranks #47 of 741 LLMs on the LLMs Tiger aggregate benchmark score (53.0), based on 14 benchmark results captured up to 2026-10-05.
- Organisation
- DeepSeek
- Type
- Open-weight / local model
- Open weights
- Yes
- Released
- 2026-07-31 (source: epoch)
- Cheapest API price
- $0.015 in / $1.28 out per million tokens (via openrouter, observed 2026-10-05T21:47Z)
- Aggregate rank
- #47 of 741 (score 53.0)
Relative to other models, DeepSeek V4 Flash 0731 (max) is strongest on GPQA Diamond (#27 of 274) and weakest on SimpleQA Verified (#44 of 66).
DeepSeek V4 Flash 0731 (max) benchmark results
| Benchmark | Score | Rank | Source | Captured |
|---|---|---|---|---|
| GPQA Diamond | 91.0% | #27 of 274 | Epoch AI Benchmarking Hub | 2026-10-05 |
| DTBench | 90.9% | #22 of 149 | Epoch AI Benchmarking Hub | 2026-10-05 |
| OTIS Mock AIME | 94.4% | #28 of 188 | Epoch AI Benchmarking Hub | 2026-10-05 |
| WeirdML | 63.0% | #17 of 112 | Epoch AI Benchmarking Hub | 2026-10-05 |
| ALE-Bench | 1,306 | #12 of 78 | Epoch AI Benchmarking Hub | 2026-10-05 |
| Chess Puzzles | 33.0% | #23 of 124 | Epoch AI Benchmarking Hub | 2026-10-05 |
| CritPt | 16.6% | #27 of 111 | Epoch AI Benchmarking Hub | 2026-10-05 |
| LMCA | 41.7% | #28 of 111 | Epoch AI Benchmarking Hub | 2026-10-05 |
| ARC-AGI-2 | 61.4% | #27 of 89 | ARC Prize official leaderboard | 2026-10-05 |
| Mystery Games | 34.0% | #18 of 57 | Epoch AI Benchmarking Hub | 2026-10-05 |
| SciCode | 49.9% | #33 of 100 | Epoch AI Benchmarking Hub | 2026-10-05 |
| FrontierMath T1-3 | 57.5% | #31 of 69 | Epoch AI Benchmarking Hub | 2026-10-05 |
| FrontierMath T4 | 24.4% | #33 of 52 | Epoch AI Benchmarking Hub | 2026-10-05 |
| SimpleQA Verified | 33.6% | #44 of 66 | Epoch AI Benchmarking Hub | 2026-10-05 |
DeepSeek V4 Flash 0731 (max) is also listed at 5 other reasoning-effort settings; this page shows the highest-effort row, which is the one the leaderboard keeps by default.
Models ranked nearby
- #44 Codex + GPT-6 Astra (53.1)
- #45 o3 + gpt-4.1 (53.0)
- #46 Claude Code + Fable 5.1 (53.0)
- #48 Qwen3.5-122B-A10B (53.0)
- #49 Claude Mythos Preview (Early) (53.0)
- #50 Solar-Open2-250B (52.9)
Compare it on the live leaderboard → · How the aggregate score works