QDNASales and integration of LLM inference and training platforms, on-premises or hybrid

Sizing: which model on which machine

This section puts 7 open models against the 8 platforms in the catalogue.

Short answer. 41 model and machine combinations produce a configuration that fits. The matrix below gives, for each, the most precise format that loads the model.

Which model fits on which machine?

Each cell gives the most precise format that loads the model on that machine, and links to the detail page with the footprint in all four formats and the remaining margin.

ModelDGX SparkMac Studio UltraDGX StationRTX PRO 6000 serverH200 SXM serverB200 SXMB300 SXMGB300 NVL72
GLM 5.2does not fitNVFP4NVFP4NVFP4FP8FP8FP16FP16
Kimi K3does not fitdoes not fitdoes not fitdoes not fitdoes not fitdoes not fitNVFP4FP16
Kimi K2.7 Codedoes not fitdoes not fitNVFP4NVFP4NVFP4FP8FP16FP16
DeepSeek V4does not fitdoes not fitdoes not fitdoes not fitNVFP4NVFP4FP8FP16
Nemotron 3 Ultradoes not fitNVFP4FP8FP8FP8FP16FP16FP16
MiniMax M3does not fitFP8FP8FP8FP16FP16FP16FP16
Qwen 3.6 27BFP16FP16FP16FP16FP16FP16FP16FP16

How to read these figures

The footprint covers the model weights. In a mixture-of-experts architecture every expert stays resident in memory even though only a fraction computes on each token: active parameters govern speed, not occupancy. The attention cache sits on top and grows with context and with the number of concurrent requests.

Method

Footprints are calculated from the parameter count and the format, with a runtime margin. They are not measured on hardware. The attention cache sits on top and depends on context and concurrency. Prices are indicative and not contractual.