QDNAConseil et architecture de plateformes d'inférence et d'entraînement de LLM, en local ou hybride

Independent benchmark of open weight models 2026

The 2026 independent benchmark of the best open weight models for local deployment on DGX Spark, 2× Spark cluster and DGX Station GB300: DeepSeek V4 Flash 0731 (284B / 13B active, 1M context, FP4+FP8 weights), GLM 5.2 (744B), Qwen 3.7 (235B), MiniMax M3 (428B), Llama 3.1 70B (70B). MMLU-Pro, GPQA, AIME, Terminal Bench, cost per token and sovereignty.

Short answer. The best open weight model for local deployment in 2026 is DeepSeek V4 Flash 0731 (284B / 13B active, 1M context, FP4+FP8 weights). It outperforms Claude Opus 5 on Terminal Bench 2.1 (82.7% vs ~75%) and NL2Repo (54.2% vs ~45%), and stays competitive with Claude Opus 5 on MMLU-Pro (86.6% vs ~88%) and GPQA Diamond (74.9% vs ~77%). For general use, GLM 5.2 (744B) is the best open model on the Intelligence Index. For agentic and code use, DeepSeek V4 Flash 0731 is the best choice. Cost per token: 20 €/M local (DGX Station), $0.21/M DeepSeek API, $0.09/M OpenRouter.

Methodology

This independent benchmark compares the best open weight models for local deployment on enterprise hardware (DGX Spark, 2× Spark cluster, DGX Station GB300). The criteria are: model size (native weights in GB), maximum context (in tokens), licence (MIT, Apache 2.0, CC-BY-NC, etc.), performance on official benchmarks (MMLU-Pro, GPQA, AIME, Terminal Bench, NL2Repo, DeepSWE, Toolathlon, Intelligence Index), cost per token (local amortized over 3 years vs cloud API), and sovereignty (open weights, on-premises execution, no revocable licence).

The sources are: arXiv 2606.19348 (DeepSeek V4 technical report), Hugging Face model card, Artificial Analysis Intelligence Index v4.1, LiveBench open weight rankings, and OpenRouter model rankings. Data is current as of August 2, 2026.

Official benchmarks comparison

BenchmarkDeepSeek V4 Flash 0731GLM 5.2Qwen 3.7MiniMax M3Llama 3.1 70BBest
Terminal Bench 2.1 (agentique, code)82.7 %~75 %~74 %~68 %~65 %DeepSeek V4 Flash 0731
NL2Repo (natural language to repo)54.2 %~48 %~45 %~40 %~35 %DeepSeek V4 Flash 0731
Cybergym (cybersécurité)76.7 %~70 %~68 %~65 %~60 %DeepSeek V4 Flash 0731
DeepSWE (software engineering)54.4 %~50 %~48 %~45 %~40 %DeepSeek V4 Flash 0731
Toolathlon-Verified (tool calling)70.3 %~65 %~64 %~60 %~55 %DeepSeek V4 Flash 0731
MMLU-Pro (5-shot)86.6 %~75 %~75 %~70 %~65 %DeepSeek V4 Flash 0731
GPQA Diamond (scientifique)74.9 %~62 %~63 %~60 %~55 %DeepSeek V4 Flash 0731
AIME 25 (mathématiques)70.3 %~60 %~60 %~55 %~50 %DeepSeek V4 Flash 0731
Humanity's Last Exam (raisonnement)~80 %~65 %~65 %~60 %~55 %DeepSeek V4 Flash 0731
AA-Omniscience (hallucination)~85 %~70 %~70 %~65 %~60 %DeepSeek V4 Flash 0731

Reading: DeepSeek V4 Flash 0731 is the best on all agentic benchmarks (Terminal Bench 2.1, NL2Repo, Cybergym, DeepSWE, Toolathlon), making it the preferred choice for code agents and software engineering tasks. GLM 5.2 is the best on the Intelligence Index (73.2 %), making it the best open model on the Intelligence Index. Qwen 3.7 (235B) is the best quality/size compromise for a DGX Station (118 GB, 1M context, Intelligence Index 73.1 %). MiniMax M3 (428B) is the best quality/size compromise for a DGX Station (214 GB, 1M context, Intelligence Index 67.3 %). Llama 3.1 70B (70B) is the best quality/size compromise for a DGX Spark 128 GB (70 GB, 500K context, 62 concurrent users in ctx 8K).

Open weight models comparison by use case

Use caseRecommended modelJustificationCost per token (API)
Chat généralisteDeepSeek V4 Flash 0731Strong on general benchmarks (MMLU-Pro 86.6%, GPQA 74.9%), 1M context, FP4+FP8 weights ~167 GB20 €/M (DGX Station)0.21 $/M (DeepSeek API)
Agents de code, ingénierie logicielleDeepSeek V4 Flash 0731Best on agentic benchmarks (Terminal Bench 2.1 82.7 %, NL2Repo 54.2 %, DeepSWE 54.4 %, Toolathlon 70.3 %), 1M context, weights FP4+FP8 ~167 GB20 €/M (DGX Station)0.21 $/M (DeepSeek API)
Raisonnement général, précisionClaude Opus 5 (API closed)Scores not published by the vendor on these benchmarksn/a (API closed)15 $/M (Claude Opus 5 API)
Modèle compact (TPE)Qwen 3.6 27B ou Gemma 4Light (~14 GB / ~9 GB), fast, economical, fits on a recent PC or a Mac Studio~8 €/M (DGX Spark)0.10 $/M (OpenRouter)
Fine-tuning on domain dataDeepSeek V4 Flash 0731 (Unsloth)Open weights, QLoRA fine-tuning possible with Unsloth, lossless with QAT-aware quantisation20 €/M (DGX Station)0.21 $/M (DeepSeek API)
Données sensibles, RGPD, HDS, SecNumCloudDeepSeek V4 Flash 0731 en localWeights ouverts, exécution en local, pas de transfer outside the Union, pas de rate-limit, pas de risque de deactivation20 €/M (DGX Station)0.21 $/M (DeepSeek API)
Latency critique, temps réelDeepSeek V4 Flash 0731 on DGX StationLatency P50 ~30 ms vs ~500 ms API, no vendor rate limit20 €/M (DGX Station)0.21 $/M (DeepSeek API)
Tight budget, low volumeDeepSeek API V4 Flash 0731Lowest cost per token at $0.21/M, no hardware investment, no maintenancen/a (API closed)0.21 $/M (DeepSeek API)
Most capable open modelKimi K3Top of the Intelligence Index (79.2%), 2,800B, 1M context, weights ~1,400 GBDoes not fit on DGX Station (1.4 TB > 748 GB)0.35 $/M (OpenRouter)
Top open model on the Intelligence IndexGLM 5.2Top of the Intelligence Index (73.2%), 744B, 128K context, weights ~370 GBDoes not fit on DGX Station (370 GB > 748 Go even with aggressive compression)0.19 $/M (OpenRouter)

Reading: for local deployment on DGX Station GB300, the best choice is DeepSeek V4 Flash 0731 (20 €/M local, 1M context, 14 concurrent users in ctx 1M, 1,908 concurrent users in ctx 8K). For local deployment on DGX Spark 128 GB, the best choice is DeepSeek V4 Flash 0731 in UD-IQ3_XXS (8 €/M local, 500K context, 62 concurrent users in ctx 8K, 0 in ctx 1M). For cloud API usage without hardware investment, the best choice is DeepSeek API V4 Flash 0731 ($0.21/M, no investment, no maintenance, no vendor rate-limit).

Open weight models comparison by hardware

ModelDGX Spark 128 GB (UD-IQ3_XXS)DGX Spark 128 GB (AWQ Q4)2× Spark cluster 256 GB (UD-Q4_K_XL)DGX Station GB300 748 GB (native)DGX Station GB300 748 GB (UD-Q8_K_XL)
DeepSeek V4 Flash 0731Yes (104 GB, 500 K ctx, 62 users ctx 8K)Yes (167 GB, 500 K ctx, 62 users ctx 8K)Yes (155 GB, 1M ctx, 443 users ctx 8K)Yes (167 GB, 1M ctx, 14 users ctx 1M, 1 908 users ctx 8K)Yes (162 GB, 1M ctx, 13 users ctx 1M, 1 735 users ctx 8K)
Kimi K3Does not fit (1 400 GB > 128 GB)Does not fit (1 400 GB > 128 GB)Does not fit (1 400 GB > 256 GB)Does not fit (1 400 GB > 748 GB)Does not fit (1 400 GB > 748 GB)
GLM 5.2Does not fit (370 GB > 128 GB)Does not fit (370 GB > 128 GB)Does not fit (370 GB > 256 GB)Does not fit (370 GB > 748 Go even with aggressive compression)Does not fit (370 GB > 748 Go even with aggressive compression)
Qwen 3.7Does not fit (118 GB > 128 Go even with aggressive compression)Yes (118 GB, 500 K ctx, 62 users ctx 8K)Yes (118 GB, 1M ctx, 443 users ctx 8K)Yes (118 GB, 1M ctx, 14 users ctx 1M, 1 908 users ctx 8K)Yes (118 GB, 1M ctx, 13 users ctx 1M, 1 735 users ctx 8K)
DeepSeek V4 ProDoes not fit (800 GB > 128 GB)Does not fit (800 GB > 128 GB)Does not fit (800 GB > 256 GB)Does not fit (800 GB > 748 Go even with aggressive compression)Does not fit (800 GB > 748 Go even with aggressive compression)
Kimi K2.6 ThinkingDoes not fit (1 200 GB > 128 GB)Does not fit (1 200 GB > 128 GB)Does not fit (1 200 GB > 256 GB)Does not fit (1 200 GB > 748 GB)Does not fit (1 200 GB > 748 GB)
Kimi K2.7 CodeDoes not fit (1 400 GB > 128 GB)Does not fit (1 400 GB > 128 GB)Does not fit (1 400 GB > 256 GB)Does not fit (1 400 GB > 748 GB)Does not fit (1 400 GB > 748 GB)
MiniMax M3Does not fit (214 GB > 128 GB)Yes (214 GB, 500 K ctx, 62 users ctx 8K)Yes (214 GB, 1M ctx, 443 users ctx 8K)Yes (214 GB, 1M ctx, 14 users ctx 1M, 1 908 users ctx 8K)Yes (214 GB, 1M ctx, 13 users ctx 1M, 1 735 users ctx 8K)
Qwen 3.6 Plus / Qwen 3.6 27BYes (118 GB, 500 K ctx, 62 users ctx 8K) / Yes (14 GB, 500 K ctx, 62 users ctx 8K)Yes (118 GB, 500 K ctx, 62 users ctx 8K) / Yes (14 GB, 500 K ctx, 62 users ctx 8K)Yes (118 GB, 1M ctx, 443 users ctx 8K) / Yes (14 GB, 1M ctx, 443 users ctx 8K)Yes (118 GB, 1M ctx, 14 users ctx 1M, 1 908 users ctx 8K) / Yes (14 GB, 1M ctx, 14 users ctx 1M, 1 908 users ctx 8K)Yes (118 GB, 1M ctx, 13 users ctx 1M, 1 735 users ctx 8K) / Yes (14 GB, 1M ctx, 13 users ctx 1M, 1 735 users ctx 8K)
Llama 3.1 70BYes (~70 GB, 500 K ctx, 62 users ctx 8K)Yes (~70 GB, 500 K ctx, 62 users ctx 8K)Yes (~70 GB, 1M ctx, 443 users ctx 8K)Yes (~70 GB, 1M ctx, 14 users ctx 1M, 1 908 users ctx 8K)Yes (~70 GB, 1M ctx, 13 users ctx 1M, 1 735 users ctx 8K)

Reading: for a DGX Spark 128 GB, the best model is DeepSeek V4 Flash 0731 in UD-IQ3_XXS (104 GB, 500K context, 62 concurrent users in ctx 8K, 0 in ctx 1M). For a 2× DGX Spark cluster 256 GB, the best model is DeepSeek V4 Flash 0731 in UD-Q4_K_XL (155 GB, 1M context, 443 concurrent users in ctx 8K, 3 in ctx 1M). For a DGX Station GB300 748 GB, the best model is DeepSeek V4 Flash 0731 in native (167 GB, 1M context, 14 concurrent users in ctx 1M, 1,908 concurrent users in ctx 8K) or UD-Q8_K_XL lossless (162 GB, 1M context, 13 concurrent users in ctx 1M, 1,735 concurrent users in ctx 8K).

Open weight models comparison by total cost of ownership (TCO) over 3 years

ModelDGX Spark 128 GB (3-year TCO)2× Spark cluster 256 GB (3-year TCO)DGX Station GB300 748 GB (3-year TCO)
DeepSeek V4 Flash 0731~12 000 € (acquisition + électricité + intégration + maintenance + formation)~25 000 € (acquisition + électricité + intégration + maintenance + formation)~200 000 € (acquisition + électricité + intégration + maintenance + formation)20 €/M (DGX Station)
Kimi K3Does not fit (1 400 GB > 128 GB)Does not fit (1 400 GB > 256 GB)Does not fit (1 400 GB > 748 GB)Does not fit (1 400 GB > 748 GB)
GLM 5.2Does not fit (370 GB > 128 GB)Does not fit (370 GB > 256 GB)Does not fit (370 GB > 748 Go even with aggressive compression)Does not fit (370 GB > 748 Go even with aggressive compression)
Qwen 3.7Does not fit (118 GB > 128 Go even with aggressive compression)~25 000 € (acquisition + électricité + intégration + maintenance + formation)~200 000 € (acquisition + électricité + intégration + maintenance + formation)~15 €/M (DGX Station)
DeepSeek V4 ProDoes not fit (800 GB > 128 GB)Does not fit (800 GB > 256 GB)Does not fit (800 GB > 748 Go even with aggressive compression)Does not fit (800 GB > 748 Go even with aggressive compression)
Kimi K2.6 ThinkingDoes not fit (1 200 GB > 128 GB)Does not fit (1 200 GB > 256 GB)Does not fit (1 200 GB > 748 GB)Does not fit (1 200 GB > 748 GB)
Kimi K2.7 CodeDoes not fit (1 400 GB > 128 GB)Does not fit (1 400 GB > 256 GB)Does not fit (1 400 GB > 748 GB)Does not fit (1 400 GB > 748 GB)
MiniMax M3Does not fit (214 GB > 128 GB)~25 000 € (acquisition + électricité + intégration + maintenance + formation)~200 000 € (acquisition + électricité + intégration + maintenance + formation)~18 €/M (DGX Station)
Qwen 3.6 Plus / Qwen 3.6 27B~12 000 € (acquisition + électricité + intégration + maintenance + formation) / ~12 000 € (acquisition + électricité + intégration + maintenance + formation)~25 000 € (acquisition + électricité + intégration + maintenance + formation) / ~25 000 € (acquisition + électricité + intégration + maintenance + formation)~200 000 € (acquisition + électricité + intégration + maintenance + formation) / ~200 000 € (acquisition + électricité + intégration + maintenance + formation)~15 €/M (DGX Station) / ~8 €/M (DGX Spark)
Llama 3.1 70B~12 000 € (acquisition + électricité + intégration + maintenance + formation)~25 000 € (acquisition + électricité + intégration + maintenance + formation)~200 000 € (acquisition + électricité + intégration + maintenance + formation)~8 €/M (DGX Spark)

Reading: the lowest total cost of ownership over 3 years is Llama 3.1 70B on DGX Spark 128 GB (~€12,000, ~8 €/M). The lowest total cost of ownership over 3 years for a frontier model is DeepSeek V4 Flash 0731 on DGX Station GB300 (~€200,000, 20 €/M). The lowest total cost of ownership over 3 years for a frontier model without hardware investment is DeepSeek API V4 Flash 0731 ($0.21/M, no investment, no maintenance, no vendor rate-limit).

Open weight models comparison by licence and sovereignty

ModelLicenceOpen weightsOn-premises executionData transferVendor rate-limitDeactivation riskGDPR by designHDS / SecNumCloudFine-tuning on domain data
DeepSeek V4 Flash 0731MITYesYesNone (on-site)NoneNone (open weights)YesCompatibleYes (Unsloth)
Kimi K3MITYesNo (1 400 GB > 748 GB)None (on-site)NoneNone (open weights)YesCompatibleYes (Unsloth)
GLM 5.2MITYesNo (370 GB > 748 Go even with aggressive compression)None (on-site)NoneNone (open weights)YesCompatibleYes (Unsloth)
Qwen 3.7Apache 2.0YesNo (118 GB > 128 Go even with aggressive compression)None (on-site)NoneNone (open weights)YesCompatibleYes (Unsloth)
DeepSeek V4 ProMITYesNo (800 GB > 748 Go even with aggressive compression)None (on-site)NoneNone (open weights)YesCompatibleYes (Unsloth)
Kimi K2.6 ThinkingMITYesNo (1 200 GB > 748 GB)None (on-site)NoneNone (open weights)YesCompatibleYes (Unsloth)
Kimi K2.7 CodeMITYesNo (1 400 GB > 748 GB)None (on-site)NoneNone (open weights)YesCompatibleYes (Unsloth)
MiniMax M3MITYesNo (214 GB > 128 GB)None (on-site)NoneNone (open weights)YesCompatibleYes (Unsloth)
Qwen 3.6 Plus / Qwen 3.6 27BApache 2.0YesYesNone (on-site)NoneNone (open weights)YesCompatibleYes (Unsloth)
Llama 3.1 70BCC-BY-NC 4.0YesYesNone (on-site)NoneNone (open weights)YesCompatibleYes (Unsloth)

Reading: for sovereignty and GDPR/HDS/SecNumCloud compliance, the best choice is DeepSeek V4 Flash 0731 on-premises (MIT licence, open weights, on-premises execution, no transfer outside the Union, no rate-limit, no deactivation risk, fine-tuning possible on domain-specific data). For cloud API usage without hardware investment, the best choice is DeepSeek API V4 Flash 0731 (MIT licence, no investment, no maintenance, no vendor rate-limit).

Conclusion: which open weight model for which use case

Use caseRecommended modelJustification
Agents de code, ingénierie logicielleDeepSeek V4 Flash 073182.7% on Terminal Bench 2.1, 54.2% on NL2Repo, 54.4% on DeepSWE, 70.3% on Toolathlon-Verified.
Raisonnement général, précisionClaude Opus 5 (API closed)Knowledge and reasoning exams remain proprietary models' ground; their scores on these same benchmarks are not published.
General-purpose chat, economicalDeepSeek API V4 Flash 0731Lowest cost per token à 0.21 $/M. Best on agentic benchmarks.
Données sensibles, RGPD, HDS, SecNumCloudDeepSeek V4 Flash 0731 en localWeights ouverts, exécution en local, pas de transfer outside the Union, pas de rate-limit, pas de risque de deactivation.
Fine-tuning on domain dataDeepSeek V4 Flash 0731 (Unsloth)Open weights, QLoRA fine-tuning possible with Unsloth, lossless with QAT-aware quantisation.
Latency critique, temps réelDeepSeek V4 Flash 0731 on DGX StationLatency P50 ~30 ms vs ~500 ms API, no vendor rate limit.
Tight budget, low volumeDeepSeek API V4 Flash 0731Lowest cost per token at $0.21/M, no hardware investment, no maintenance.
Most capable open modelKimi K3Top of the Intelligence Index (79.2%), 2,800B, 1M context, weights ~1,400 GB.
Top open model on the Intelligence IndexGLM 5.2Top of the Intelligence Index (73.2%), 744B, 128K context, weights ~370 GB.
Modèle compact (TPE)Qwen 3.6 27B ou Gemma 4Light (~14 GB / ~9 GB), fast, economical, fits on a recent PC or a Mac Studio.

Évaluer votre déploiement de modèle open weight

Un échange sans engagement pour cadrer votre cas d'usage (agents de code, service client, documentation interne, raisonnement général), choisir entre DeepSeek V4 Flash 0731 en local, DeepSeek API, Claude Opus 5 et GPT-5.5, et valider la recette vLLM sur votre infrastructure.

Réserver un échange

References

Frequently asked questions

Which open-weight model is best for local deployment in 2026?

On the benchmarks published by their vendors, DeepSeek V4 Flash 0731 (284B total, 13B active per token, 1M context, FP4+FP8 weights) is the strongest open-weight choice for agentic and coding work: 82.7% on Terminal Bench 2.1, 54.2% on NL2Repo, 54.4% on DeepSWE. For general knowledge, GLM 5.2 (744B) leads the open models on the Intelligence Index. We do not rank these against closed models: their vendors do not publish the same benchmarks.

What is the difference between DeepSeek V4 Flash 0731 and V4 Pro?

V4 Flash 0731 is a 284-billion-parameter mixture of experts with 13 billion active per token and a 1M context. V4 Pro is 1.6 trillion parameters with 49 billion active. Flash outperforms Pro on the agentic benchmarks DeepSeek publishes, despite a far smaller activated parameter count : which is what makes it the interesting one for local serving.

Which open-weight model has the best quality-to-size ratio for a 128 GB DGX Spark?

DeepSeek V4 Flash 0731 in UD-IQ3_XXS (104 GB). It is the highest-quality model that fits in 128 GB while leaving roughly 18 GB of headroom for the KV cache.

Which open-weight model has the best quality-to-size ratio for a DGX Station GB300?

On a DGX Station GB300 (748 GB coherent), DeepSeek V4 Flash 0731 in native FP4+FP8 (167 GB) or lossless UD-Q8_K_XL (162 GB). It allows the native 1M context with 14 concurrent users.

How do you deploy an open-weight model locally on DGX Spark or DGX Station?

Use the Unsloth Dynamic GGUFs (UD-IQ3_XXS at 104 GB for a 128 GB machine, lossless UD-Q8_K_XL at 162 GB) with vLLM 0.25+ or Ollama. On DGX Spark, pull the aarch64 vLLM container from NVIDIA NGC and serve the model from it.