QDNASales and integration of LLM inference and training platforms, on-premises or hybrid

B300 SXM

HGX B300, Blackwell Ultra 8 GPU

B300 SXM
Memory288 GB HBM3e per GPU (2304 GB)
Bandwidth8 TB/s per GPU
Computeapproximately 15 PFLOPS FP4 per GPU (120 total)
NVLinkNVLink 5, 1.8 TB/s
Power drawapproximately 1.4 kW per GPU, liquid cooling
Targetfrontier models, long context

What the machine runs

Inference and training of large models at scale. The largest open-weight models served with headroom, including mixture-of-experts and a one-million-token context. Recommended runtime: vLLM with KV cache offload.

Positioning

The high-end server adds fifty percent more memory per GPU over the B200, for reasoning and long-context workloads. An eight-GPU configuration falls between 430,000 and 550,000 euros, as an estimate.

Manufacturers

HGX B300 examples: Dell PowerEdge XE9712, Supermicro SYS-821GB-TNRX, HPE Cray XD670, Lenovo ThinkSystem SR680b, Gigabyte G383-P00. ASUS rounds out the offering with the ESC series.

GPUDirect Storage

On platforms fitted with ConnectX cards, GPUDirect Storage technology establishes a direct path between NVMe or NVMe over Fabric storage and GPU memory, bypassing the CPU buffer. Sustained throughput reaches approximately 50 gigabytes per second per link. Compatible storage arrays, such as NetApp, VAST, DDN or WEKA, feed training and large-scale RAG at full speed.

Support

Open stack supported by QDNA and the communities. NVIDIA AI Enterprise option (NIM, NeMo, Triton) with a service-level commitment, per card.

Configure this hardware

A no-obligation call to size your platform.

Book a call