QDNASales and integration of LLM inference and training platforms, on-premises or hybrid

GB300 NVL72

Rack scale, 72 liquid-cooled GPUs

GB300 NVL72
GPU72 Blackwell Ultra and 36 Grace (2,592 cores)
Memoryapproximately 20.7 TB HBM3e (40 TB of fast memory)
Computeapproximately 1.1 ExaFLOPS FP4 dense
NVLink130 TB/s across a 72-GPU domain
Power drawapproximately 120 kW, liquid cooling
Form factorfull 42U rack

What the machine runs

A single 72-chip supercomputer serves frontier models and a private AI cloud. Recommended runtime: disaggregated vLLM with LiteLLM.

Positioning

Large enterprises host their own frontier model in a single rack, in a sovereign AI cloud.

Manufacturers

Rack integration from Dell, HPE, Lenovo, Supermicro, ASUS and Gigabyte, with liquid cooling.

GPUDirect Storage

On platforms equipped with ConnectX cards, GPUDirect Storage technology establishes a direct path between NVMe or NVMe over Fabric storage and GPU memory, bypassing the CPU buffer. Sustained throughput reaches approximately 50 gigabytes per second per link. Compatible storage arrays, such as NetApp, VAST, DDN or WEKA, feed training and large-scale RAG at full speed.

Support

Open stack supported by QDNA and the communities. NVIDIA AI Enterprise option (NIM, NeMo, Triton) with a service-level agreement, per card.

Configure this hardware

A no-obligation call to size your platform.

Book a call