QDNASales and integration of LLM inference and training platforms, on-premises or hybrid

RTX PRO 6000 server

Air-cooled GPU server, Blackwell Server Edition

RTX PRO 6000 server
Memory96 GB GDDR7 per GPU (2 to 8 GPUs, up to 768 GB)
Computearound 4 PFLOPS FP4 with sparsity per GPU
Bandwidth1.6 TB/s, PCIe Gen5
IsolationMIG up to four 24 GB instances
Power draw600 W per GPU, air-cooled
Density2 to 8 GPUs per server

What the machine runs

Several models from 70 to 200 billion parameters in parallel, RAG and multi-GPU fine-tuning. Recommended runtime: vLLM. Backbone of the dwarfstar stack.

Positioning

The sovereign workhorse runs air-cooled on standard servers, with the best cost-per-token ratio.

Manufacturers

Certified servers from Dell, HPE, Lenovo, Supermicro, ASUS and Gigabyte.

GPUDirect Storage

On platforms fitted with ConnectX cards, GPUDirect Storage technology establishes a direct path between NVMe or NVMe over Fabric storage and GPU memory, bypassing the CPU buffer. Sustained throughput reaches around 50 gigabytes per second per link. Compatible arrays, such as NetApp, VAST, DDN or WEKA, feed training and large-scale RAG at full speed.

Support

Open stack supported by QDNA and the communities. Optional NVIDIA AI Enterprise (NIM, NeMo, Triton) with a per-card service-level agreement.

Configure this hardware

A no-obligation call to size your platform.

Book a call