QDNASales and integration of LLM inference and training platforms, on-premises or hybrid

B200 SXM

HGX B200, 8-GPU Blackwell

B200 SXM
Memory180 GB HBM3e per GPU (1440 GB total)
Bandwidth7.7 TB/s per GPU
Computeabout 9 PFLOPS dense FP4 per GPU (72 total)
NVLink1.8 TB/s NVLink 5 (14.4 TB/s total)
Powerabout 1 kW per GPU
Transistors208 billion per GPU

What the machine runs

Frontier-level inference and training, high FP4 throughput. Recommended runtime: vLLM in NVFP4 format.

Positioning

The move to Blackwell reaches fifteen times the Hopper generation on the largest models.

Manufacturers

HGX B200 systems from Dell, HPE, Lenovo, Supermicro, ASUS and Gigabyte.

GPUDirect Storage

On platforms fitted with ConnectX cards, GPUDirect Storage technology establishes a direct path between NVMe or NVMe over Fabric storage and GPU memory, bypassing the CPU buffer. Sustained throughput reaches about 50 gigabytes per second per link. Compatible arrays, such as NetApp, VAST, DDN or WEKA, feed training and massive-scale RAG at full speed.

Support

Open stack supported by QDNA and the communities. NVIDIA AI Enterprise option (NIM, NeMo, Triton) with a per-card service-level agreement.

Configure this hardware

A no-obligation call to size your platform.

Book a call