Best Open-Source and Open-Weight LLMs (2026)
Kimi K3 (max) leads this ranking with a score of 55.7 (50 is the neutral midpoint). This page ranks models on every benchmark we track, restricted to models whose weights are public. Data last captured 2026-10-05.
| # | Model | Organisation | Score | Overall rank | Cheapest input $/M |
|---|---|---|---|---|---|
| 1 | Kimi K3 (max) | Moonshot | 55.7 | #17 | $1.29 |
| 2 | DeepSeek V4 Pro 0813 (max) | DeepSeek | 54.3 | #25 | $0.660 |
| 3 | GLM-5.3 (max) | Z.ai (Zhipu AI) | 53.8 | #31 | $0.070 |
| 4 | DeepSeek R1 (0528) | deepseek-ai | 53.6 | #35 | $0.400 |
| 5 | NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 | nvidia | 53.1 | #42 | โ |
| 6 | NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | nvidia | 53.1 | #43 | โ |
| 7 | DeepSeek V4 Flash 0731 (max) | DeepSeek | 53.0 | #47 | $0.015 |
| 8 | Qwen3.5-122B-A10B | Qwen | 53.0 | #48 | $0.113 |
| 9 | Solar-Open2-250B | upstage | 52.9 | #50 | โ |
| 10 | Kimi K2.5 (Fireworks) | Moonshot | 52.8 | #53 | $0.450 |
| 11 | Qwen3-Coder-30B-A3B-Instruct | Qwen | 52.7 | #56 | $0.070 |
| 12 | Step-3.5-Flash | stepfun-ai | 52.5 | #66 | $0.100 |
| 13 | DeepSeek-V3.2-Exp (high) | DeepSeek | 52.4 | #68 | $0.270 |
| 14 | Darwin-180B-RSI | FINAL-Bench | 52.3 | #71 | โ |
| 15 | Intern-S2-Preview | internlm | 52.3 | #72 | โ |
| 16 | Darwin-397B-ZTC | FINAL-Bench | 52.2 | #75 | โ |
| 17 | Ornith-1.5-397B | ornith-ai | 52.2 | #78 | โ |
| 18 | Qwen3.8-2.4T-A95B | Qwen | 52.1 | #79 | $2.00 |
| 19 | EXAONE-4.5-33B | LGAI-EXAONE | 52.1 | #80 | โ |
| 20 | NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | nvidia | 52.0 | #82 | โ |
| 21 | K-EXAONE-236B-A23B | LGAI-EXAONE | 52.0 | #83 | โ |
| 22 | Darwin-398B-JGOS | FINAL-Bench | 52.0 | #87 | โ |
| 23 | Qwen3.8-Flash-Next | Qwen | 51.9 | #88 | โ |
| 24 | Hy4-preview | tencent | 51.9 | #90 | $0.751 |
| 25 | Hy3 | tencent | 51.9 | #91 | $0.083 |
| 26 | ollama/qwen2.5-coder:32b | Alibaba | 51.9 | #94 | โ |
| 27 | Kimi K2 | Moonshot AI | 51.8 | #98 | $0.550 |
| 28 | Ornith-1.5-35B-A3B | ornith-ai | 51.8 | #100 | โ |
| 29 | DeepSeek-V2.5-1210 | DeepSeek | 51.7 | #101 | โ |
| 30 | Darwin-36B-Opus | FINAL-Bench | 51.7 | #104 | โ |
How this ranking is built
Rows are the highest-effort variant of each model, ranked by the all-benchmark aggregate described in the methodology; open-weight status comes from the model's published licence.