QDNAAdvisory and architecture for LLM inference and training platforms, on-premises or hybrid

B300 SXM

HGX B300, Blackwell Ultra 8 GPU

B300 SXM
Indicative priceabout €430k to €690k excl. VAT: 8-GPU HGX B300 server BIZON X9000 G5 listed at $496,728 excl. VAT on 2 September 2026 (November 2026 delivery) and Supermicro AS-8126GS-NB3RT listed at $795,000 excl. VAT at Vipera the same day, converted at the ECB rate of $1.159 per €1 (price guide)
Memory288 GB HBM3e per GPU (2,304 GB; NVIDIA rounds to 2.1 TB)
Bandwidth8 TB/s per GPU
Compute13.5 PFLOPS dense FP4 per GPU (108 total on the HGX B300 board, 144 with sparsity)
NVLinkNVLink 5, 1.8 TB/s
Power drawup to 1,400 W per GPU, liquid cooling
Targetfrontier models, long context

What the machine runs

Inference and training of large models at scale. The largest open-weight models served with headroom, including mixture-of-experts and a one-million-token context. Recommended runtime: vLLM with KV cache offload.

Positioning

The high-end server adds 60% more memory per GPU over the B200 (288 GB versus 180 GB), for reasoning and long-context workloads. An eight-GPU configuration is listed from $496,728 to $795,000 excl. VAT at BIZON and Vipera on 2 September 2026, i.e. about €430k to €690k excl. VAT at the ECB rate.

Manufacturers

HGX B300 examples: Dell PowerEdge XE9712, Supermicro SYS-821GB-TNRX, HPE Cray XD670, Lenovo ThinkSystem SR680b, Gigabyte G383-P00. ASUS rounds out the offering with the ESC series.

GPUDirect Storage

On platforms fitted with ConnectX cards, GPUDirect Storage technology establishes a direct path between NVMe or NVMe over Fabric storage and GPU memory, bypassing the CPU buffer. Compatible storage arrays, such as NetApp, VAST, DDN or WEKA, feed training and large-scale RAG at full speed.

Support

Open stack supported by QDNA and the communities. NVIDIA AI Enterprise option (NIM, NeMo, Triton) with a service-level agreement, licensed per GPU.

Configure this hardware

A no-obligation call to size your platform.

Book a call

Which models fit on B300 SXM?

The machine offers 2,304 GB of memory. 7 of the 7 open models in the catalogue load on it, in the format shown.

ModelBillion parametersMost precise format that fitsSizing page
GLM 5.2744FP16GLM 5.2 on B300 SXM
Kimi K32,800NVFP4Kimi K3 on B300 SXM
Kimi K2.7 Code1,000FP16Kimi K2.7 Code on B300 SXM
DeepSeek V41,600FP8DeepSeek V4 on B300 SXM
Nemotron 3 Ultra550FP16Nemotron 3 Ultra on B300 SXM
MiniMax M3428FP16MiniMax M3 on B300 SXM
Qwen 3.8 27B27FP16Qwen 3.8 27B on B300 SXM

See the full sizing matrix.

Frequently asked questions

How much does a configured b300 sxm cost for an LLM?

Pricing depends on configuration and supplier. The page shows a dated public price; we provide a detailed quote after a scoping call.

Which LLM fits in a b300 sxm?

Depends on the quantization format and the input context. The table on the page indicates the recommended maximum model size.

Which runtimes support the b300 sxm?

vLLM, llama.cpp, Triton Inference Server, depending on the chosen framework. The model, format and GPU combinations measured by QDNA are published in the measurements section.