QDNAAdvisory and architecture for LLM inference and training platforms, on-premises or hybrid

B200 SXM

HGX B200, 8-GPU Blackwell

B200 SXM
Indicative priceon quotation, depending on the configuration (price guide)
Memory180 GB HBM3e per GPU (1,440 GB total; NVIDIA rounds to 1.4 TB)
Bandwidth7.7 TB/s per GPU
Compute9 PFLOPS dense FP4 per GPU (72 total, 144 with sparsity)
NVLink1.8 TB/s NVLink 5 (14.4 TB/s total)
Power1,000 W per GPU (TGP)
Transistors208 billion per GPU

What the machine runs

Frontier-level inference and training, high FP4 throughput. Recommended runtime: vLLM in NVFP4 format.

Positioning

According to NVIDIA, the move to Blackwell reaches up to fifteen times the Hopper generation in inference on the largest models.

Manufacturers

HGX B200 systems from Dell, HPE, Lenovo, Supermicro, ASUS and Gigabyte.

GPUDirect Storage

On platforms fitted with ConnectX cards, GPUDirect Storage technology establishes a direct path between NVMe or NVMe over Fabric storage and GPU memory, bypassing the CPU buffer. Compatible arrays, such as NetApp, VAST, DDN or WEKA, feed training and massive-scale RAG at full speed.

Support

Open stack supported by QDNA and the communities. NVIDIA AI Enterprise option (NIM, NeMo, Triton) with a service-level agreement, licensed per GPU.

Configure this hardware

A no-obligation call to size your platform.

Book a call

Which models fit on B200 SXM?

The machine offers 1,440 GB of memory. 6 of the 7 open models in the catalogue load on it, in the format shown.

ModelBillion parametersMost precise format that fitsSizing page
GLM 5.2744FP8GLM 5.2 on B200 SXM
Kimi K2.7 Code1,000FP8Kimi K2.7 Code on B200 SXM
DeepSeek V41,600NVFP4DeepSeek V4 on B200 SXM
Nemotron 3 Ultra550FP16Nemotron 3 Ultra on B200 SXM
MiniMax M3428FP16MiniMax M3 on B200 SXM
Qwen 3.8 27B27FP16Qwen 3.8 27B on B200 SXM

See the full sizing matrix.

Frequently asked questions

How much does a configured b200 sxm cost for an LLM?

Pricing depends on configuration and supplier. No public price is listed for this configuration; we provide a detailed quote after a scoping call.

Which LLM fits in a b200 sxm?

Depends on the quantization format and the input context. The table on the page indicates the recommended maximum model size.

Which runtimes support the b200 sxm?

vLLM, llama.cpp, Triton Inference Server, depending on the chosen framework. The model, format and GPU combinations measured by QDNA are published in the measurements section.