Qwen3.8 Max (xhigh): benchmark scores, price and ranking

Qwen3.8 Max (xhigh) from Alibaba ranks #39 of 741 LLMs on the LLMs Tiger aggregate benchmark score (53.2), based on 12 benchmark results captured up to 2026-10-05.

Organisation
Alibaba
Type
Frontier model
Open weights
No
Released
2026-08-02 (source: epoch)
Cheapest API price
$1.65 in / $4.95 out per million tokens (via deepinfra, observed 2026-10-05T21:47Z)
Aggregate rank
#39 of 741 (score 53.2)

Relative to other models, Qwen3.8 Max (xhigh) is strongest on OTIS Mock AIME (#11 of 188) and weakest on FrontierSWE (#14 of 16).

Qwen3.8 Max (xhigh) benchmark results

Qwen3.8 Max (xhigh) benchmark scores, ranked against all models measured on each benchmark
BenchmarkScoreRankSourceCaptured
OTIS Mock AIME99.4%#11 of 188Epoch AI Benchmarking Hub2026-10-05
GPQA Diamond92.7%#18 of 274Epoch AI Benchmarking Hub2026-10-05
DTBench92.0%#19 of 149Epoch AI Benchmarking Hub2026-10-05
LMCA46.2%#21 of 111Epoch AI Benchmarking Hub2026-10-05
Mystery Games38.0%#13 of 57Epoch AI Benchmarking Hub2026-10-05
Chess Puzzles29.0%#29 of 124Epoch AI Benchmarking Hub2026-10-05
FrontierMath T1-374.7%#17 of 69Epoch AI Benchmarking Hub2026-10-05
FrontierMath T446.3%#20 of 52Epoch AI Benchmarking Hub2026-10-05
SimpleQA Verified45.8%#28 of 66Epoch AI Benchmarking Hub2026-10-05
GDP.pdf23.2%#10 of 22Epoch AI Benchmarking Hub2026-10-05
DeepSWE57.5%#14 of 19Epoch AI Benchmarking Hub2026-10-05
FrontierSWE15.8%#14 of 16Epoch AI Benchmarking Hub2026-10-05

Qwen3.8 Max (xhigh) is also listed at 2 other reasoning-effort settings; this page shows the highest-effort row, which is the one the leaderboard keeps by default.

Models ranked nearby

Compare it on the live leaderboard → · How the aggregate score works