QDNASales and integration of LLM inference and training platforms, on-premises or hybrid

H200 SXM server

HGX H200, 8 SXM GPUs

H200 SXM server
Memory141 GB HBM3e per GPU, 8 GPUs (1128 GB)
Bandwidth4.8 TB/s per GPU
NVLink900 GB/s per GPU, full NVSwitch mesh
ComputeFP8 and FP16, tensor parallelism across 8 GPUs
Powerup to 700 W per GPU
Form factor8-GPU HGX board, 8U server

What the machine runs

A 1128 GB pool for inference and training at scale: fine-tuning, LoRA and continued pre-training across eight GPUs meshed via NVLink. Recommended runtime: vLLM with tensor parallelism.

Positioning

The HGX H200 board makes eight GPUs work as a single accelerator via NVSwitch, for training and inference of large models. NVIDIA AI Enterprise licensing is sold per GPU.

Vendors

HGX H200 systems from Dell, HPE, Lenovo, Supermicro, ASUS and Gigabyte.

GPUDirect Storage

On platforms fitted with ConnectX cards, GPUDirect Storage technology establishes a direct path between NVMe or NVMe over Fabric storage and GPU memory, bypassing the CPU buffer. Sustained throughput reaches around 50 gigabytes per second per link. Compatible storage arrays, such as NetApp, VAST, DDN or WEKA, feed training and massive RAG at full speed.

Support

Open stack supported by QDNA and the communities. Optional NVIDIA AI Enterprise (NIM, NeMo, Triton) with a service-level commitment, per GPU.

Configure this hardware

A no-obligation call to size your platform.

Book a call