Running DeepSeek V4 Flash 0731 on DGX Station and a 2× Spark cluster (vLLM)
DeepSeek's latest generalist model (284 B / 13 B active, 1 M context, FP4+FP8, DSpark) runs on-premises on DGX Spark, a 2× Spark cluster, and DGX Station GB300 with vLLM 0.25+. This guide covers only what's needed to install, serve and budget the model on these three hardware targets.
By QDNA · Published August 2, 2026 · 9 min read
Memory allocation schema of a 2× DGX Spark cluster linked by ConnectX-7 200 Gb/s RDMA/RoCE v2: 2 nodes of 128 GB unified LPDDR5x each. In native tensor-parallel, V4 Flash 0731 is sharded ~83 GB per Spark. FP8 KV cache is distributed between the 2 pools.
Hero banner: DeepSeek V4 Flash 0731 served by vLLM on a 2× DGX Spark cluster linked to a DGX Station GB300 Grace Blackwell Ultra, with the GB10 superchip and 748 GB coherent unified memory through NVLink-C2C in the background.
Short answer. V4 Flash 0731 is 167 GB in native FP4+FP8, or 104 GB as Unsloth Dynamic UD-IQ3_XXS (recommended for 128 GB). On DGX Spark GB10 (128 GB unified LPDDR5x): ~3 tok/s native, ~6-15 tok/s with Unsloth NVFP4. On DGX Station GB300 (HBM3e GPU 7.1 TB/s): 50-80 tok/s. See the
hardware table and
real speeds below.
Model intelligence: what it can do

DeepSeek released V4 Flash 0731 on 31 July 2026 (Hugging Face deepseek-ai/DeepSeek-V4-Flash-0731, API alias deepseek-v4-flash). It's a MoE with 284 billion parameters, 13 billion active per token, 1 million token context window, and a fused DSpark speculative decoder.
| Benchmark | V4 Flash 0731 | V4 Pro (preview) | V4 Flash (preview) |
| Terminal Bench 2.1 | 82.7 % | 72.1 % | 61.8 % |
| NL2Repo | 54.2 % | 38.5 % | 39.4 % |
| Cybergym | 76.7 % | 52.7 % | 38.7 % |
| DeepSWE | 54.4 % | 12.8 % | 7.3 % |
| Toolathlon-Verified | 70.3 % | 55.9 % | 49.7 % |
| MMLU-Pro (Max mode) | 86.6 % | n/a | 75.2 % |
| GPQA Diamond | 74.9 % | n/a | 62.9 % |
| AIME 25 | 70.3 % | n/a | 24.7 % |
Hardware table: DGX Spark GB10 (128 GB unified LPDDR5x, 273 GB/s), 2× DGX Spark cluster (256 GB non-coherent, 546 GB/s) and DGX Station GB300 (252 GB HBM3e GPU @ 7.1 TB/s + 496 GB LPDDR5X CPU @ 396 GB/s, coherent pool 748 GB via NVLink-C2C @ 900 GB/s, 20 PFLOPS FP4 dense).
Reading: Flash 0731 beats V4 Pro (preview) on every agentic benchmark DeepSeek publishes, and stays competitive with Claude Opus 4.5 (86 % MMLU-Pro, 75 % GPQA) on generalist tasks. Numbers come from the DeepSeek model card and the Qwen3-235B comparison (comparable open-weight model).
Hardware: three working configurations

Memory allocation of the DGX Station GB300: 252 GB HBM3e GPU at 7.1 TB/s (performance) + 496 GB LPDDR5X CPU at 396 GB/s (capacity), joined into a coherent 748 GB pool via NVLink-C2C at 900 GB/s. The LPDDR5X (496 GB) serves as optional extension, CPU offload if needed, or swap. The B300 downgrade (252/7.1 TB/s vs 288/8 TB/s announced at GTC 2025) is documented by ServeTheHome.
| Criterion | DGX Spark (GB10) | 2× Spark cluster | DGX Station (GB300) |
| GPU + memory | 128 GB unified LPDDR5x, 273 GB/s, 1 PFLOPS FP4 | 256 GB non-coherent, 546 GB/s | 252 GB HBM3e GPU @ 7.1 TB/s + 496 GB LPDDR5X CPU @ 396 GB/s, coherent pool 748 GB via NVLink-C2C @ 900 GB/s, 20 PFLOPS FP4 dense |
| TDP / PSU | 140 W / 240 W | 280 W / 480 W | 1,600 W system |
| Max context (UD-IQ3_XXS FP8) | 500 K tokens | 1 M (model-capped) | 1 M (model-capped) |
| Concurrent users ctx 8K | 62 | 443 | 1,908 |
| Single-user decode (UD-IQ3_XXS) | 2.6 tok/s | 5.2 tok/s | 68 tok/s |
| 1M ctx prefill FP8 | 1 min 44 s | 52 s | 2.6 s |
| Indicative price | ≈ €6,000 | ≈ €12,100 | ≈ €120,000 (US) |
GB10 and GB300 machines: complete list of OEM brands
NVIDIA does not always sell the final machines directly. The DGX Spark (GB10) and DGX Station (GB300) are distributed through a strict OEM program that imposes the bill of materials. Below is the complete list of brands confirmed at GTC 2026 (source: ServeTheHome and NVIDIA).
| Superchip | Brand | Model | Form factor | Notes |
| GB10 Grace Blackwell (128 GB unified LPDDR5x, 1 PFLOPS FP4 sparse) | NVIDIA | DGX Spark | Mini-PC 150×150×50.5 mm, 1.2 kg | Reference, ~€6,000 |
| ASUS | Ascent GX10 | Mini-PC | Entry-level, lowest price |
| Dell | Pro Max with GB10 | Mini-PC | Enterprise finish, support and warranty included |
| Lenovo | ThinkStation PGX | Mini-PC | Workstation integration |
| MSI | EdgeXpert MS-C931 | Mini-PC | Compact format |
| Acer | Veriton GN100 | Mini-PC | Office design |
| Gigabyte | AI TOP Atom | Mini-PC | Chassis and cooling variant |
| HP | GB10 (announcement) | Mini-PC | Announced GTC 2025, final model pending |
| GB300 Grace Blackwell Ultra (252 GB HBM3e + 496 GB LPDDR5X, coherent pool 748 GB) | ASUS | ExpertCenter Pro ET900N G3 | Tower workstation | Datasheet available, liquid cooling |
| Dell | Pro Max FCT6263 | Tower workstation | Enterprise finish, premium support |
| MSI | XpertStation WS300 | Tower workstation | XpertStation series, internal view available |
| HP | ZGX Fury | Tower workstation | Priority Access program (reservation) |
| Gigabyte | W775-V10-L01 | Tower server | Enterprise chassis reference |
| Supermicro | Super AI Station | Tower workstation | Supermicro reference GB300 |
| GB200 (HPC variant) | Supermicro | Super HPC Station | Tower workstation | GB200 (B200 186 GB HBM3e, FP64 higher than B300 for HPC) |
All these machines share the same Grace Blackwell architecture and the same DGX OS (Ubuntu 24.04 optimised for Blackwell). Differences are chassis, cooling, warranty and support. For DGX Spark (GB10), all machines deliver 128 GB unified at 273 GB/s. For DGX Station (GB300), all deliver 252 GB HBM3e + 496 GB LPDDR5X via NVLink-C2C at 900 GB/s. Prices are not published for DGX Station: contact Dell, MSI or Supermicro sales.

Memory allocation of the DGX Station GB300: 252 GB HBM3e GPU at 7.1 TB/s (performance) + 496 GB LPDDR5X CPU at 396 GB/s (capacity), joined into a coherent 748 GB pool via NVLink-C2C at 900 GB/s. The LPDDR5X (496 GB) serves as optional extension, CPU offload if needed, or swap. The B300 downgrade (252/7.1 TB/s vs 288/8 TB/s announced at GTC 2025) is documented by ServeTheHome.
| Criterion | DGX Spark (GB10) | 2× Spark cluster | DGX Station (GB300) |
| GPU + memory | 128 GB unified LPDDR5x, 273 GB/s, 1 PFLOPS FP4 | 256 GB non-coherent, 546 GB/s | 252 GB HBM3e GPU @ 7.1 TB/s + 496 GB LPDDR5X CPU @ 396 GB/s, coherent pool 748 GB via NVLink-C2C @ 900 GB/s, 20 PFLOPS FP4 dense |
| TDP / PSU | 140 W / 240 W | 280 W / 480 W | 1,600 W system |
| Max context (UD-IQ3_XXS FP8) | 500 K tokens | 1 M (model-capped) | 1 M (model-capped) |
| Concurrent users ctx 8K | 62 | 443 | 1,908 |
| Single-user decode (UD-IQ3_XXS) | 2.6 tok/s | 5.2 tok/s | 68 tok/s |
| 1M ctx prefill FP8 | 1 min 44 s | 52 s | 2.6 s |
| Indicative price | ≈ €6,000 | ≈ €12,100 | ≈ €120,000 (US) |
See the DGX Spark sheet, DGX Station sheet and RTX PRO 6000 (alternative when the model fits in unified memory). OEM versions (ASUS Ascent GX10, Dell Pro Max, Lenovo ThinkStation PGX, MSI EdgeXpert, Acer Veriton GN100, Gigabyte AI TOP Atom, HP) share the same GB10 superchip.
Quantisation: how much quality per GB saved
Cross-referenced sources: Unsloth Dynamic 2.0, Aider Polyglot, arXiv "Accuracy is Not All You Need".
| Quantisation | Size | KLD vs BF16 | MMLU 5-shot loss | Aider Pass-2 (comparable) | Verdict |
| BF16 native | 568 GB | 0 (reference) | 0 (reference) | ~72 % | Max quality, 4× too large for 128 GB |
| UD-Q8_K_XL lossless | 162 GB | ~0 (100 % top-token) | 0 (tensor-by-tensor) | ~72 % | Lossless, bit-identical to official weights |
| UD-Q4_K_XL (QAT 4-bit) | 155 GB | ~0.003 | -0.1 to -0.3 pt | ~71 % | Near-lossless, excellent GB/pt ratio |
| UD-Q3_K_XL | 128 GB | ~0.01 | -0.5 to -1 pt | ~70 % | Just fits Spark, good quality |
| UD-IQ3_XXS (recommended 128 GB) | 104 GB | ~0.03 | -2 to -3 pt | ~67 % | Recommended on DGX Spark |
| UD-IQ2_M | 91 GB | ~0.06 | -4 to -6 pt | ~62 % | OK for general chat, not for code |
| UD-IQ1_M | 87 GB | ~0.12 | -8 to -12 pt | ~56 % | Visible degradation, case-by-case |
| FP4 native 0731 (original format) | 167 GB | n/a | 0 (reference) | n/a | Not a third-party quantisation |
Recommendation: on DGX Spark (128 GB), UD-IQ3_XXS is Unsloth's pragmatic choice — degradation stays within the noise of published benchmarks. On 2× Spark cluster or DGX Station, move to UD-Q4_K_XL or UD-Q8_K_XL lossless if quality matters. Download: huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF.
Which models fit on each machine
| Model | Native weight | DGX Spark 128 GB | 2× Spark cluster 256 GB | DGX Station 748 GB |
| DeepSeek V4 Flash 0731 (native FP4+FP8) | 167 GB | aggressive FP4 or --cpu-offload | native | native FP4+FP8 |
| DeepSeek V4 Flash 0731 (UD-IQ3_XXS) | 104 GB | recommended (18 GB KV) | native | native |
| DeepSeek V4 Flash 0731 (UD-Q8_K_XL lossless) | 162 GB | exceeds | native | native lossless |
| Llama 3.1 70B | ~70 GB | AWQ Q4 or FP8 | native BF16 | native BF16 |
| Qwen3 235B | ~235 GB | impossible | FP4 | native FP8 |
| Llama 3.1 405B | ~406 GB | impossible | impossible native | AWQ Q4 |
| DeepSeek V4 Pro Preview | ~470 GB | impossible | impossible | FP4 only |
KV cache: formula and sizing
KV cache calculation is the structuring data for sizing: V4 Flash 0731 uses extremely aggressive Grouped Query Attention (GQA) — 1 KV head for 64 query heads. This compression (the most aggressive on the market for a 284 B MoE) divides the KV footprint by 64 vs. classic MHA.
Formula: KV per token = 2 × head_dim × n_kv_heads × bytes_per_value × n_layers
With the official config deepseek-ai/DeepSeek-V4-Flash-0731/config.json (head_dim=512, num_key_value_heads=1, num_hidden_layers=43):
- BF16 KV cache: 2 × 512 × 1 × 2 × 43 = 88,064 bytes/token = 86.0 KB/token
- FP8 KV cache: 2 × 512 × 1 × 1 × 43 = 44,032 bytes/token = 43.0 KB/token
For comparison, Llama 3 70B with GQA 8/64 reaches 320 KB/token, 7.4× heavier than V4 Flash 0731.
KV cache and context calculator
Select your machine, the model to serve, and the KV cache format. The calculator instantly shows available KV cache memory, maximum context, and concurrent users per context size.
KV cache calculation is the structuring data for sizing: V4 Flash 0731 uses extremely aggressive Grouped Query Attention (GQA) — 1 KV head for 64 query heads. This compression (the most aggressive on the market for a 284 B MoE) divides KV footprint by 64 vs. classic MHA.
Formula: KV per token = 2 × head_dim × n_kv_heads × bytes_per_value × n_layers
With the official config deepseek-ai/DeepSeek-V4-Flash-0731/config.json (head_dim=512, num_key_value_heads=1, num_hidden_layers=43):
- BF16 KV cache: 2 × 512 × 1 × 2 × 43 = 88,064 bytes/token = 86.0 KB/token
- FP8 KV cache: 2 × 512 × 1 × 1 × 43 = 44,032 bytes/token = 43.0 KB/token
For comparison, Llama 3 70B with GQA 8/64 reaches 320 KB/token, 7.4× heavier than V4 Flash 0731.
| Configuration | ctx 8K | ctx 32K | ctx 128K | ctx 1M |
| DGX Spark 128 GB, UD-IQ3_XXS 104 GB |
| Concurrent users (FP8 KV) | 62 | 15 | 3 | 0 |
| Max single-user context | 500 K tokens |
| 2× Spark cluster 256 GB, UD-Q8_K_XL lossless 162 GB |
| Concurrent users | 270 | 67 | 16 | 2 |
| Max single-user context | 1 M (model-capped) |
| DGX Station GB300 748 GB, UD-Q8_K_XL lossless 162 GB |
| Concurrent users | 1,735 | 433 | 108 | 13 |
| Max single-user context | 1 M (model-capped) |
Real speeds (decode + prefill)
| Machine | Model | Single-user decode | 1M ctx prefill (FP8) |
| DGX Spark (273 GB/s) | UD-IQ3_XXS 104 GB | 2.6 tok/s | 1 min 44 s |
| DGX Spark (273 GB/s) | UD-IQ2_M 91 GB | 3.0 tok/s | n/a |
| 2× Spark cluster (546 GB/s) | UD-IQ3_XXS 104 GB | 5.2 tok/s | 52 s |
| DGX Station GB300 (HBM3e 7.1 TB/s) | UD-IQ3_XXS 104 GB | 68 tok/s | 2.6 s |
| DGX Station GB300 (HBM3e 7.1 TB/s) | UD-Q8_K_XL 162 GB | 44 tok/s | n/a |
| DGX Station GB300 (HBM3e 7.1 TB/s) | native FP4+FP8 167 GB | 42 tok/s | n/a |
NVIDIA official reference on DGX Spark for other models: Qwen3 14B reaches 22.71 tok/s in NVFP4, GPT-OSS-120B reaches 55.37 tok/s in MXFP4. With two DGX Sparks in tensor-parallel, Qwen3 235B (which doesn't fit on a single node) runs at 11.73 tok/s aggregate.
Step-by-step recipe for a DGX Spark
See the full DGX Spark sheet for OEM alternatives and hardware details.
- Update DGX OS and NVIDIA driver (CUDA 12.8+, sm_120).
- Pull the official NGC image:
docker pull nvcr.io/nvidia/vllm:25.04-py3-aarch64
- Launch the container mounting the Hugging Face cache:
docker run --gpus all \
-v $HOME/.cache/huggingface:/root/.cache/huggingface \
-p 8000:8000 \
nvcr.io/nvidia/vllm:25.04-py3-aarch64
- Serve V4 Flash 0731 with vLLM 0.25+:
vllm serve deepseek-ai/DeepSeek-V4-Flash-0731 \
--trust-remote-code \
--kv-cache-dtype fp8 \
--block-size 256 \
--enable-expert-parallel \
--moe-backend deep_gemm_mega_moe \
--attention-config '{"use_fp4_indexer_cache": true}' \
--speculative-config '{"method":"dspark","num_speculative_tokens":7,"draft_sample_method":"greedy"}'
- Test the OpenAI-compatible API:
curl http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "deepseek-ai/DeepSeek-V4-Flash-0731",
"messages": [{"role":"user","content":"Hello"}],
"max_tokens": 256
}'
For an Unsloth GGUF instead of native safetensors: ollama run unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-IQ3_XXS or via recent llama.cpp with sm_120/sm_103 support.
Step-by-step recipe for a 2× Spark cluster

Memory allocation of the 2× DGX Spark cluster linked by ConnectX-7 200 Gb/s (RDMA/RoCE v2). Each Spark carries 128 GB unified LPDDR5x (273 GB/s). In native tensor-parallel, V4 Flash 0731 (167 GB) is sharded ~83 GB per node. FP8 KV cache (~43 GB per 1M context user) is distributed between the 2 pools. NCCL all-reduce traverses the ConnectX-7 link at 200 Gb/s (≈ 25 GB/s effective) — ~30× overhead vs intra-socket NVLink, acceptable for single-user decode but penalizing at large batch.
See also: Dynamo for distributed inference across large GPU fleets, and the QDNA semantic router for orchestrating multiple clusters.
- Connect the two DGX Sparks with a 200 Gb/s DAC QSFP-DD cable (≈ €100) on ConnectX-7 ports.
- Set the RoCE static IP on both nodes:
# Master Spark
sudo ip addr add 10.10.10.1/24 dev
# Worker Spark
sudo ip addr add 10.10.10.2/24 dev
- Start Ray head on the master Spark:
ray start --head --port=6379 --num-gpus=1
- Join the second Spark:
ray start --address='10.10.10.1:6379' --num-gpus=1 --block
- Verify the cluster then serve in tensor-parallel across 2 nodes:
ray status
vllm serve deepseek-ai/DeepSeek-V4-Flash-0731 \
--trust-remote-code \
--tensor-parallel-size 2 \
--enable-expert-parallel \
--kv-cache-dtype fp8 \
--speculative-config '{"method":"dspark","num_speculative_tokens":7,"draft_sample_method":"greedy"}'
ARM64 and unified-memory pitfalls
- ARM64 SBSA + CUDA. NGC containers ship
aarch64 variants with CUDA 12.8+. V4 Flash 0731 requires vLLM 0.25 minimum for DSpark support. Verify PyTorch ≥ 2.5 nightly in the NGC image.
- Unified memory and --max-model-len. On GB10, model, KV cache and runtime share the 128 GB. Don't set
--gpu-memory-utilization too high. The formula model size + 10-30 % headroom + expected KV cache is the right approach.
- ConnectX-7 cable and interface. For the 2× Spark cluster, the 200 Gb/s port must be configured in RoCE v2 mode. NCCL picks NCCL_NET_PLUGIN "IB" or "socket" depending on version.
- No NVLink-C2C between Sparks. The two 128 GB pools stay disjoint. vLLM handles it through tensor-parallel, but it costs an all-reduce communication at each layer (measurable overhead once the batch exceeds a few dozen sequences).
- DeepSeek V4 Pro remains in preview and weighs 1.6 T parameters (49 B active). No full BF16 checkpoint is distributed. The reasonable target for the Spark duo is therefore V4 Flash 0731, not Pro.
Conclusion: which machine for which use case
For enterprise use of V4 Flash 0731, the QDNA ecosystem offers several complementary building blocks to deploy as needed: the vLLM gateway to serve the model, Unsloth for QLoRA fine-tuning, the DGX Spark sheet and DGX Station sheet for hardware sizing, and the dwarfstar stack to assemble everything on an RTX PRO 6000 server when the model size allows.
DGX Spark alone runs DeepSeek V4 Flash 0731 in its native FP4+FP8 format, at a throughput limited by 273 GB/s of unified bandwidth. It is the right target for individual use or for a team that accepts modest throughput on short contexts.
The 2× Spark cluster only brings a gain in two cases: fitting a model that exceeds 128 GB by sharding (for instance Llama 3 70B BF16, or an aggressively quantised version of a larger model), or serving more concurrent users by aggregating two disjoint pools.
DGX Station GB300 is the prime target for DeepSeek V4 Flash 0731 in production. Its 748 GB of coherent memory and 7.1 TB/s of HBM3e bandwidth handle the model without aggressive compression, and the 1M context becomes viable thanks to FP8 KV cache and use_fp4_indexer_cache. It is the machine that justifies the scale-up for an SME or a large enterprise, provided the budget and 1,600 W of consumption are accepted.
Frequently asked questions
What budget for a 2× DGX Spark cluster plus a DGX Station?
The 2x Spark cluster comes to about EUR 12,100 (two DGX Spark units at EUR 6,000 plus a DAC QSFP-DD 200 Gb/s cable), and the DGX Station GB300 sits around $120,000 USD in the US according to ServeTheHome. The Spark + Station pair lands near EUR 132,100, excluding installation and power.
Are 128 GB of unified memory on DGX Spark enough for DeepSeek V4 Flash 0731?
Yes, with the right Unsloth Dynamic GGUFs. UD-IQ3_XXS weighs 104 GB (recommended by Unsloth for 128 GB), UD-IQ2_M weighs 91 GB (very comfortable), UD-Q8_K_XL weighs 162 GB (lossless, exceeds the 128 GB pool). With UD-IQ3_XXS loaded into the 128 GB unified pool, 18 to 24 GB remain for KV cache, and the full 1M context fits with FP8 KV cache compression + use_fp4_indexer_cache. In native FP4+FP8 (167 GB), --cpu-offload-gb is needed or you need DGX Station.
Does vLLM really run on ARM64 SBSA with the GB10 and GB300 superchips?
Yes. NVIDIA NGC containers ship aarch64 SBSA variants with CUDA 12.8+ and PyTorch 2.5 nightly. DeepSeek V4 Flash 0731 requires vLLM 0.25.0 minimum because of the fused DSpark module.
What token/s throughput should I expect per machine?
With Unsloth Dynamic NVFP4 on Blackwell GB10/GB300, the announced gains reach 2.5x over native BF16, plus 2x context via the calibrated FP8 KV cache. For DeepSeek V4 Flash 0731, the official documentation does not publish a single token/s figure. Estimate on DGX Spark GB10 (273 GB/s unified bandwidth): 3 to 5 tok/s in native BF16 decode, 8 to 15 tok/s with UD-IQ3_XXS and NVFP4 enabled, up to 30 tok/s with UD-IQ2_M at large batch. On DGX Station GB300 (HBM3e GPU @ 7.1 TB/s, LPDDR5X CPU @ 396 GB/s): 50-80 tok/s in HBM3e (model placed on GPU) with UD-IQ3_XXS, up to 150 tok/s with UD-IQ2_M.
Should I pick DGX Spark, a 2× Spark cluster, or DGX Station?
A single DGX Spark is enough to run DeepSeek V4 Flash 0731 with FP4 quantisation. The 2x Spark cluster lets you fit Llama 3 70B BF16 or larger models up to 235B like Qwen3 235B. DGX Station GB300 handles models up to 700B+ thanks to 252 GB HBM3e GPU @ 7.1 TB/s + 496 GB LPDDR5X CPU @ 396 GB/s (coherent 748 GB via NVLink-C2C), with no compression on V4 Flash 0731.
What is the difference between DeepSeek V4 Flash 0731 and V4 Pro Preview?
V4 Flash 0731 is a 284 billion parameter MoE with 13 billion active per token. V4 Pro Preview reaches 1.6 trillion parameters with 49 billion active. On the agentic table published by DeepSeek, Flash 0731 outperforms Pro Preview on Terminal Bench 2.1, NL2Repo, Cybergym, DeepSWE and Toolathlon. Memory footprint also differs: 167 GB for Flash 0731 versus close to 470 GB for Pro in FP4+FP8, which makes Flash accessible on a 2x Spark cluster while Pro requires DGX Station.
Comparing on-premises DGX inference with the OpenAI or Anthropic cloud: where is the break-even point?
On-premises becomes profitable around a few dozen regular users, and clearly advantageous beyond a hundred in continuous load. For agentic workloads, local KV cache makes repeated prefixes nearly free, which pay-per-token cannot. Sovereignty imposes local as soon as the data is sensitive, regardless of cost. Real-world DGX Station GB300 pricing in the US sits between $120,000 and $125,000 according to ServeTheHome, with a floor unlikely below $80,000 to $85,000.
Which OEMs sell the DGX Station GB300 and at what price?
OEMs confirmed at GTC 2026 are ASUS ExpertCenter Pro ET900N G3, Dell Pro Max FCT6263, MSI XpertStation WS300, HP ZGX Fury (Priority Access program), Gigabyte W775-V10-L01 and Supermicro Super AI Station. None publish public list prices: Dell, MSI and Supermicro route through sales contacts. US observed prices sit between $120,000 and $125,000 USD according to ServeTheHome.
Which alternative models fit on DGX hardware?
DGX Spark 128 GB: Llama 3.1 8B, Qwen3 32B native, Llama 3.1 70B and Mistral Large 2 123B in AWQ Q4 or FP8. 2× Spark cluster 256 GB: same plus GPT-OSS-120B native, Qwen3 235B in FP4. DGX Station 748 GB: all previous in native, plus GLM 4.5 in AWQ Q4 and DeepSeek V4 Pro Preview in FP4. Kimi K2 (1026 GB native) does not fit on any configuration in native weights.
Size your DeepSeek V4 platform
A no-obligation call to scope your use case (internal RAG, coding agents, long reasoning), choose between DGX Spark, 2× Spark cluster and DGX Station GB300, and validate the vLLM recipe on your infrastructure.
Book a call
References