Open weight models for local deployment 2026: complete comparison
The 2026 comparison of the best open weight models for local deployment on DGX Spark, 2× Spark cluster and DGX Station GB300: DeepSeek V4 Flash 0731, Kimi K3, GLM 5.2, Qwen 3.7, MiniMax M3, Llama 3.1 70B. Size, context, licence, performance, cost and sovereignty.
Methodology
This comparison evaluates the best open weight models for local deployment on enterprise hardware (DGX Spark, 2× Spark cluster, DGX Station GB300). The criteria are: model size (native weights in GB), maximum context (in tokens), licence (MIT, Apache 2.0, CC-BY-NC, etc.), performance on official benchmarks (MMLU-Pro, GPQA, AIME, Terminal Bench, NL2Repo, DeepSWE, Toolathlon, Intelligence Index), cost per token (local amortized over 3 years vs cloud API), and sovereignty (open weights, on-premises execution, no revocable licence).
The sources are: arXiv 2606.19348 (DeepSeek V4 technical report), Hugging Face model card, Artificial Analysis Intelligence Index v4.1, LiveBench open weight rankings, and OpenRouter model rankings. Data is current as of August 2, 2026.
Open weight models comparison for local deployment
| Model | Parameters | Active per token | Context | Native size | Licence | Intelligence Index score | Cost per token (local) | Cost per token (API) |
|---|---|---|---|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | 284 B | 13 B | 1 M | 167 GB | MIT | 82.7 % (Terminal Bench 2.1) | 20 €/M (DGX Station) | 0.21 $/M (DeepSeek API) |
| Kimi K3 | 2 800 B | ~50 B | 1 M | ~1 400 GB | MIT | 79.2 % (Intelligence Index) | Does not fit on DGX Station (1.4 TB > 748 GB) | 0.35 $/M (OpenRouter) |
| GLM 5.2 | 744 B | ~20 B | 128 K | ~370 GB | MIT | 73.2 % (Intelligence Index) | Does not fit on DGX Station (370 GB > 748 Go even with aggressive compression) | 0.19 $/M (OpenRouter) |
| Qwen 3.7 Flash / Qwen 3.7 Max | 235 B | ~15 B | 1 M | ~118 GB | Apache 2.0 | 73.1 % (Intelligence Index) | ~15 €/M (DGX Station) | 0.18 $/M (OpenRouter) |
| DeepSeek V4 Pro | 1 600 B | 49 B | 1 M | ~800 GB | MIT | 71.6 % (Intelligence Index) | Does not fit on DGX Station (800 GB > 748 Go even with aggressive compression) | 0.05 $/M (OpenRouter) |
| Kimi K2.6 Thinking | 2 400 B | ~40 B | 1 M | ~1 200 GB | MIT | 70.5 % (Intelligence Index) | Does not fit on DGX Station (1.2 TB > 748 GB) | 0.17 $/M (OpenRouter) |
| Kimi K2.7 Code | 2 800 B | ~45 B | 1 M | ~1 400 GB | MIT | 68.4 % (Intelligence Index) | Does not fit on DGX Station (1.4 TB > 748 GB) | 0.10 $/M (OpenRouter) |
| MiniMax M3 | 428 B | ~15 B | 1 M | ~214 GB | MIT | 67.3 % (Intelligence Index) | ~18 €/M (DGX Station) | 0.06 $/M (OpenRouter) |
| Qwen 3.6 Plus / Qwen 3.6 27B | 235 B / 27 B | ~15 B / ~10 B | 1 M / 128 K | ~118 Go / ~14 GB | Apache 2.0 | 68.9 % (Intelligence Index) | ~15 €/M (DGX Station) / ~8 €/M (DGX Spark) | 0.23 $/M (OpenRouter) |
| Llama 3.1 70B | 70 B | ~10 B | 128 K | ~70 GB | CC-BY-NC 4.0 | ~65 % (estimation) | ~8 €/M (DGX Spark) | 0.10 $/M (OpenRouter) |
Reading: DeepSeek V4 Flash 0731 offers the best quality-to-size ratio for local deployment in 2026 : 284B total, 13B active per token, 1M context, FP4+FP8 weights at roughly 167 GB, 82.7% on Terminal Bench 2.1. Kimi K3 (2,800B) tops the Intelligence Index among open models but cannot be deployed on a DGX Station (1.4 TB against 748 GB). GLM 5.2 (744B) also leads the index but does not fit either without aggressive compression (370 GB against 748 GB). Qwen 3.7 Flash / Qwen 3.7 Max (235B) is the best quality-to-size compromise for a DGX Station (118 GB, 1M context, Intelligence Index 73.1%).
Open weight models comparison by use case
| Use case | Recommended model | Justification | Cost per token (local) | Cost per token (API) |
|---|---|---|---|---|
| Chat généraliste | DeepSeek V4 Flash 0731 | Strong on general benchmarks (MMLU-Pro 86.6%, GPQA 74.9%), 1M context, FP4+FP8 weights ~167 GB | 20 €/M (DGX Station) | 0.21 $/M (DeepSeek API) |
| Agents de code, ingénierie logicielle | DeepSeek V4 Flash 0731 | Best on agentic benchmarks (Terminal Bench 2.1 82.7 %, NL2Repo 54.2 %, DeepSWE 54.4 %, Toolathlon 70.3 %), 1M context, weights FP4+FP8 ~167 GB | 20 €/M (DGX Station) | 0.21 $/M (DeepSeek API) |
| Raisonnement général, précision | Claude Opus 5 (API closed) | Scores not published by the vendor on these benchmarks | n/a (API closed) | 15 $/M (Claude Opus 5 API) |
| Modèle compact (TPE) | Qwen 3.6 27B ou Gemma 4 | Light (~14 GB / ~9 GB), fast, economical, fits on a recent PC or a Mac Studio | ~8 €/M (DGX Spark) | 0.10 $/M (OpenRouter) |
| Fine-tuning on domain data | DeepSeek V4 Flash 0731 (Unsloth) | Open weights, QLoRA fine-tuning possible with Unsloth, lossless with QAT-aware quantisation | 20 €/M (DGX Station) | 0.21 $/M (DeepSeek API) |
| Données sensibles, RGPD, HDS, SecNumCloud | DeepSeek V4 Flash 0731 en local | Weights ouverts, exécution en local, pas de transfer outside the Union, pas de rate-limit, pas de risque de deactivation | 20 €/M (DGX Station) | 0.21 $/M (DeepSeek API) |
| Latency critique, temps réel | DeepSeek V4 Flash 0731 on DGX Station | Latency P50 ~30 ms vs ~500 ms API, no vendor rate limit | 20 €/M (DGX Station) | 0.21 $/M (DeepSeek API) |
| Tight budget, low volume | DeepSeek API V4 Flash 0731 | Lowest cost per token at $0.21/M, no hardware investment, no maintenance | n/a (API closed) | 0.21 $/M (DeepSeek API) |
| Most capable open model | Kimi K3 | Top of the Intelligence Index (79.2%), 2,800B, 1M context, weights ~1,400 GB | Does not fit on DGX Station (1.4 TB > 748 GB) | 0.35 $/M (OpenRouter) |
| Top open model on the Intelligence Index | GLM 5.2 | Top of the Intelligence Index (73.2%), 744B, 128K context, weights ~370 GB | Does not fit on DGX Station (370 GB > 748 Go even with aggressive compression) | 0.19 $/M (OpenRouter) |
Reading: for local deployment on DGX Station GB300, the best choice is DeepSeek V4 Flash 0731 (20 €/M local, 1M context, 14 concurrent users in ctx 1M, 1,908 concurrent users in ctx 8K). For local deployment on DGX Spark 128 GB, the best choice is DeepSeek V4 Flash 0731 in UD-IQ3_XXS (8 €/M local, 500K context, 62 concurrent users in ctx 8K, 0 in ctx 1M). For cloud API usage without hardware investment, the best choice is DeepSeek API V4 Flash 0731 ($0.21/M, no investment, no maintenance, no vendor rate-limit).
Open weight models comparison by hardware
| Model | DGX Spark 128 GB (UD-IQ3_XXS) | DGX Spark 128 GB (AWQ Q4) | 2× Spark cluster 256 GB (UD-Q4_K_XL) | DGX Station GB300 748 GB (native) | DGX Station GB300 748 GB (UD-Q8_K_XL) |
|---|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | Yes (104 GB, 500 K ctx, 62 users ctx 8K) | Yes (167 GB, 500 K ctx, 62 users ctx 8K) | Yes (155 GB, 1M ctx, 443 users ctx 8K) | Yes (167 GB, 1M ctx, 14 users ctx 1M, 1 908 users ctx 8K) | Yes (162 GB, 1M ctx, 13 users ctx 1M, 1 735 users ctx 8K) |
| Kimi K3 | Does not fit (1 400 GB > 128 GB) | Does not fit (1 400 GB > 128 GB) | Does not fit (1 400 GB > 256 GB) | Does not fit (1 400 GB > 748 GB) | Does not fit (1 400 GB > 748 GB) |
| GLM 5.2 | Does not fit (370 GB > 128 GB) | Does not fit (370 GB > 128 GB) | Does not fit (370 GB > 256 GB) | Does not fit (370 GB > 748 Go even with aggressive compression) | Does not fit (370 GB > 748 Go even with aggressive compression) |
| Qwen 3.7 Flash / Qwen 3.7 Max | Does not fit (118 GB > 128 Go even with aggressive compression) | Yes (118 GB, 500 K ctx, 62 users ctx 8K) | Yes (118 GB, 1M ctx, 443 users ctx 8K) | Yes (118 GB, 1M ctx, 14 users ctx 1M, 1 908 users ctx 8K) | Yes (118 GB, 1M ctx, 13 users ctx 1M, 1 735 users ctx 8K) |
| DeepSeek V4 Pro | Does not fit (800 GB > 128 GB) | Does not fit (800 GB > 128 GB) | Does not fit (800 GB > 256 GB) | Does not fit (800 GB > 748 Go even with aggressive compression) | Does not fit (800 GB > 748 Go even with aggressive compression) |
| Kimi K2.6 Thinking | Does not fit (1 200 GB > 128 GB) | Does not fit (1 200 GB > 128 GB) | Does not fit (1 200 GB > 256 GB) | Does not fit (1 200 GB > 748 GB) | Does not fit (1 200 GB > 748 GB) |
| Kimi K2.7 Code | Does not fit (1 400 GB > 128 GB) | Does not fit (1 400 GB > 128 GB) | Does not fit (1 400 GB > 256 GB) | Does not fit (1 400 GB > 748 GB) | Does not fit (1 400 GB > 748 GB) |
| MiniMax M3 | Does not fit (214 GB > 128 GB) | Yes (214 GB, 500 K ctx, 62 users ctx 8K) | Yes (214 GB, 1M ctx, 443 users ctx 8K) | Yes (214 GB, 1M ctx, 14 users ctx 1M, 1 908 users ctx 8K) | Yes (214 GB, 1M ctx, 13 users ctx 1M, 1 735 users ctx 8K) |
| Qwen 3.6 Plus / Qwen 3.6 27B | Yes (118 GB, 500 K ctx, 62 users ctx 8K) / Yes (14 GB, 500 K ctx, 62 users ctx 8K) | Yes (118 GB, 500 K ctx, 62 users ctx 8K) / Yes (14 GB, 500 K ctx, 62 users ctx 8K) | Yes (118 GB, 1M ctx, 443 users ctx 8K) / Yes (14 GB, 1M ctx, 443 users ctx 8K) | Yes (118 GB, 1M ctx, 14 users ctx 1M, 1 908 users ctx 8K) / Yes (14 GB, 1M ctx, 14 users ctx 1M, 1 908 users ctx 8K) | Yes (118 GB, 1M ctx, 13 users ctx 1M, 1 735 users ctx 8K) / Yes (14 GB, 1M ctx, 13 users ctx 1M, 1 735 users ctx 8K) |
| Llama 3.1 70B | Yes (~70 GB, 500 K ctx, 62 users ctx 8K) | Yes (~70 GB, 500 K ctx, 62 users ctx 8K) | Yes (~70 GB, 1M ctx, 443 users ctx 8K) | Yes (~70 GB, 1M ctx, 14 users ctx 1M, 1 908 users ctx 8K) | Yes (~70 GB, 1M ctx, 13 users ctx 1M, 1 735 users ctx 8K) |
Reading: for a DGX Spark 128 GB, the best model is DeepSeek V4 Flash 0731 in UD-IQ3_XXS (104 GB, 500K context, 62 concurrent users in ctx 8K, 0 in ctx 1M). For a 2× DGX Spark cluster 256 GB, the best model is DeepSeek V4 Flash 0731 in UD-Q4_K_XL (155 GB, 1M context, 443 concurrent users in ctx 8K, 3 in ctx 1M). For a DGX Station GB300 748 GB, the best model is DeepSeek V4 Flash 0731 in native (167 GB, 1M context, 14 concurrent users in ctx 1M, 1,908 concurrent users in ctx 8K) or UD-Q8_K_XL lossless (162 GB, 1M context, 13 concurrent users in ctx 1M, 1,735 concurrent users in ctx 8K).
Open weight models comparison by total cost of ownership (TCO) over 3 years
| Model | DGX Spark 128 GB (3-year TCO) | 2× Spark cluster 256 GB (3-year TCO) | DGX Station GB300 748 GB (3-year TCO) | Cost per token (3-year TCO) |
|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | ~12 000 € (acquisition + électricité + intégration + maintenance + formation) | ~25 000 € (acquisition + électricité + intégration + maintenance + formation) | ~200 000 € (acquisition + électricité + intégration + maintenance + formation) | 20 €/M (DGX Station) |
| Kimi K3 | Does not fit (1 400 GB > 128 GB) | Does not fit (1 400 GB > 256 GB) | Does not fit (1 400 GB > 748 GB) | Does not fit (1 400 GB > 748 GB) |
| GLM 5.2 | Does not fit (370 GB > 128 GB) | Does not fit (370 GB > 256 GB) | Does not fit (370 GB > 748 Go even with aggressive compression) | Does not fit (370 GB > 748 Go even with aggressive compression) |
| Qwen 3.7 Flash / Qwen 3.7 Max | Does not fit (118 GB > 128 Go even with aggressive compression) | ~25 000 € (acquisition + électricité + intégration + maintenance + formation) | ~200 000 € (acquisition + électricité + intégration + maintenance + formation) | ~15 €/M (DGX Station) |
| DeepSeek V4 Pro | Does not fit (800 GB > 128 GB) | Does not fit (800 GB > 256 GB) | Does not fit (800 GB > 748 Go even with aggressive compression) | Does not fit (800 GB > 748 Go even with aggressive compression) |
| Kimi K2.6 Thinking | Does not fit (1 200 GB > 128 GB) | Does not fit (1 200 GB > 256 GB) | Does not fit (1 200 GB > 748 GB) | Does not fit (1 200 GB > 748 GB) |
| Kimi K2.7 Code | Does not fit (1 400 GB > 128 GB) | Does not fit (1 400 GB > 256 GB) | Does not fit (1 400 GB > 748 GB) | Does not fit (1 400 GB > 748 GB) |
| MiniMax M3 | Does not fit (214 GB > 128 GB) | ~25 000 € (acquisition + électricité + intégration + maintenance + formation) | ~200 000 € (acquisition + électricité + intégration + maintenance + formation) | ~18 €/M (DGX Station) |
| Qwen 3.6 Plus / Qwen 3.6 27B | ~12 000 € (acquisition + électricité + intégration + maintenance + formation) / ~12 000 € (acquisition + électricité + intégration + maintenance + formation) | ~25 000 € (acquisition + électricité + intégration + maintenance + formation) / ~25 000 € (acquisition + électricité + intégration + maintenance + formation) | ~200 000 € (acquisition + électricité + intégration + maintenance + formation) / ~200 000 € (acquisition + électricité + intégration + maintenance + formation) | ~15 €/M (DGX Station) / ~8 €/M (DGX Spark) |
| Llama 3.1 70B | ~12 000 € (acquisition + électricité + intégration + maintenance + formation) | ~25 000 € (acquisition + électricité + intégration + maintenance + formation) | ~200 000 € (acquisition + électricité + intégration + maintenance + formation) | ~8 €/M (DGX Spark) |
Reading: the lowest total cost of ownership over 3 years is Llama 3.1 70B on DGX Spark 128 GB (~€12,000, ~8 €/M). The lowest total cost of ownership over 3 years for a frontier model is DeepSeek V4 Flash 0731 on DGX Station GB300 (~€200,000, 20 €/M). The lowest total cost of ownership over 3 years for a frontier model without hardware investment is DeepSeek API V4 Flash 0731 ($0.21/M, no investment, no maintenance, no vendor rate-limit).
Open weight models comparison by licence and sovereignty
| Model | Licence | Open weights | On-premises execution | Data transfer | Vendor rate-limit | Deactivation risk | GDPR by design | HDS / SecNumCloud | Fine-tuning on domain data |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | MIT | Yes | Yes | None (on-site) | None | None (open weights) | Yes | Compatible | Yes (Unsloth) |
| Kimi K3 | MIT | Yes | No (1 400 GB > 748 GB) | None (on-site) | None | None (open weights) | Yes | Compatible | Yes (Unsloth) |
| GLM 5.2 | MIT | Yes | No (370 GB > 748 Go even with aggressive compression) | None (on-site) | None | None (open weights) | Yes | Compatible | Yes (Unsloth) |
| Qwen 3.7 Flash / Qwen 3.7 Max | Apache 2.0 | Yes | No (118 GB > 128 Go even with aggressive compression) | None (on-site) | None | None (open weights) | Yes | Compatible | Yes (Unsloth) |
| DeepSeek V4 Pro | MIT | Yes | No (800 GB > 748 Go even with aggressive compression) | None (on-site) | None | None (open weights) | Yes | Compatible | Yes (Unsloth) |
| Kimi K2.6 Thinking | MIT | Yes | No (1 200 GB > 748 GB) | None (on-site) | None | None (open weights) | Yes | Compatible | Yes (Unsloth) |
| Kimi K2.7 Code | MIT | Yes | No (1 400 GB > 748 GB) | None (on-site) | None | None (open weights) | Yes | Compatible | Yes (Unsloth) |
| MiniMax M3 | MIT | Yes | No (214 GB > 128 GB) | None (on-site) | None | None (open weights) | Yes | Compatible | Yes (Unsloth) |
| Qwen 3.6 Plus / Qwen 3.6 27B | Apache 2.0 | Yes | Yes | None (on-site) | None | None (open weights) | Yes | Compatible | Yes (Unsloth) |
| Llama 3.1 70B | CC-BY-NC 4.0 | Yes | Yes | None (on-site) | None | None (open weights) | Yes | Compatible | Yes (Unsloth) |
Reading: for sovereignty and GDPR/HDS/SecNumCloud compliance, the best choice is DeepSeek V4 Flash 0731 on-premises (MIT licence, open weights, on-premises execution, no transfer outside the Union, no rate-limit, no deactivation risk, fine-tuning possible on domain-specific data). For cloud API usage without hardware investment, the best choice is DeepSeek API V4 Flash 0731 (MIT licence, no investment, no maintenance, no vendor rate-limit).
Conclusion: which open weight model for which use case
| Use case | Recommended model | Justification |
|---|---|---|
| Agents de code, ingénierie logicielle | DeepSeek V4 Flash 0731 | 82.7% on Terminal Bench 2.1, 54.2% on NL2Repo, 54.4% on DeepSWE, 70.3% on Toolathlon-Verified. |
| Raisonnement général, précision | Claude Opus 5 (API closed) | Knowledge and reasoning exams remain proprietary models' ground; their scores on these same benchmarks are not published. |
| General-purpose chat, economical | DeepSeek API V4 Flash 0731 | Lowest cost per token à 0.21 $/M. Best on agentic benchmarks. |
| Données sensibles, RGPD, HDS, SecNumCloud | DeepSeek V4 Flash 0731 en local | Weights ouverts, exécution en local, pas de transfer outside the Union, pas de rate-limit, pas de risque de deactivation. |
| Fine-tuning on domain data | DeepSeek V4 Flash 0731 (Unsloth) | Open weights, QLoRA fine-tuning possible with Unsloth, lossless with QAT-aware quantisation. |
| Latency critique, temps réel | DeepSeek V4 Flash 0731 on DGX Station | Latency P50 ~30 ms vs ~500 ms API, no vendor rate limit. |
| Tight budget, low volume | DeepSeek API V4 Flash 0731 | Lowest cost per token at $0.21/M, no hardware investment, no maintenance. |
| Most capable open model | Kimi K3 | Top of the Intelligence Index (79.2%), 2,800B, 1M context, weights ~1,400 GB. |
| Top open model on the Intelligence Index | GLM 5.2 | Top of the Intelligence Index (73.2%), 744B, 128K context, weights ~370 GB. |
| Modèle compact (TPE) | Qwen 3.6 27B ou Gemma 4 | Light (~14 GB / ~9 GB), fast, economical, fits on a recent PC or a Mac Studio. |
Évaluer votre déploiement de modèle open weight
Un échange sans engagement pour cadrer votre cas d'usage (agents de code, service client, documentation interne, raisonnement général), choisir entre DeepSeek V4 Flash 0731 en local, DeepSeek API, Claude Opus 5 et GPT-5.5, et valider la recette vLLM sur votre infrastructure.
Réserver un échangeReferences
- DeepSeek V4 Flash 0731, dépôt officiel Hugging Face
- Rapport technique DeepSeek V4 (arXiv 2606.19348)
- Artificial Analysis Intelligence Index v4.1
- LiveBench open weight rankings
- OpenRouter model rankings
- Unsloth, documentation DeepSeek V4
- Fiche QDNA : DeepSeek V4
- Fiche QDNA : vLLM
- Fiche QDNA : DGX Station
- Fiche QDNA : DGX Spark
Frequently asked questions
Which open-weight model is best for local deployment in 2026?
On the benchmarks published by their vendors, DeepSeek V4 Flash 0731 (284B total, 13B active per token, 1M context, FP4+FP8 weights) is the strongest open-weight choice for agentic and coding work: 82.7% on Terminal Bench 2.1, 54.2% on NL2Repo, 54.4% on DeepSWE. For general knowledge, GLM 5.2 (744B) leads the open models on the Intelligence Index. We do not rank these against closed models: their vendors do not publish the same benchmarks.
What is the difference between DeepSeek V4 Flash 0731 and V4 Pro?
V4 Flash 0731 is a 284-billion-parameter mixture of experts with 13 billion active per token and a 1M context. V4 Pro is 1.6 trillion parameters with 49 billion active. Flash outperforms Pro on the agentic benchmarks DeepSeek publishes, despite a far smaller activated parameter count : which is what makes it the interesting one for local serving.
Which open-weight model has the best quality-to-size ratio for a 128 GB DGX Spark?
DeepSeek V4 Flash 0731 in UD-IQ3_XXS (104 GB). It is the highest-quality model that fits in 128 GB while leaving roughly 18 GB of headroom for the KV cache.
Which open-weight model has the best quality-to-size ratio for a DGX Station GB300?
On a DGX Station GB300 (748 GB coherent), DeepSeek V4 Flash 0731 in native FP4+FP8 (167 GB) or lossless UD-Q8_K_XL (162 GB). It allows the native 1M context with 14 concurrent users.
How do you deploy an open-weight model locally on DGX Spark or DGX Station?
Use the Unsloth Dynamic GGUFs (UD-IQ3_XXS at 104 GB for a 128 GB machine, lossless UD-Q8_K_XL at 162 GB) with vLLM 0.25+ or Ollama. On DGX Spark, pull the aarch64 vLLM container from NVIDIA NGC and serve the model from it.