QDNASales and integration of LLM inference and training platforms, on-premises or hybrid

DGX Spark

AI workstation, GB10 superchip

DGX Spark
Memory128 GB unified LPDDR5X
Compute1 PFLOP FP4 (1000 TOPS)
ChipGB10 Grace Blackwell, 20 Arm cores
NetworkConnectX-7 200 GbE, cluster of 1 to 4 nodes
Power draw170 W, wall outlet
Form factor150 mm desktop

What the machine runs

Models up to 200 billion parameters quantised on a single node. Directly connecting two nodes lets the platform address up to 405 billion parameters, and a ConnectX-7 switch extends the cluster further. Recommended runtime: llama.cpp or vLLM locally.

Positioning

A sovereign workstation for a small business puts a private model on the desk, with no reliance on the cloud.

Manufacturers

NVIDIA offers the DGX Spark through its partners, including ASUS, Dell, Gigabyte, HP, Lenovo and MSI.

GPUDirect Storage

On platforms fitted with ConnectX cards, GPUDirect Storage technology establishes a direct path between NVMe or NVMe over Fabric storage and GPU memory, bypassing the CPU buffer. Sustained throughput reaches around 50 gigabytes per second per link. Compatible storage arrays, such as NetApp, VAST, DDN or WEKA, feed training and large-scale RAG at full speed.

Support

Open stack supported by QDNA and the communities. Optional NVIDIA AI Enterprise (NIM, NeMo, Triton) with a per-card service-level agreement.

Configure this hardware

A no-commitment conversation to size your platform.

Book a call