QDNAAdvisory and architecture for LLM inference and training platforms, on-premises or hybrid

Serving a model: which engine for which model

The engine decides throughput, concurrency and the formats accepted.

Short answer. 27 engine and model combinations are documented. Each page gives the minimum memory and the platforms that meet it.

By engine

Dynamo

Dynamo targets distributed serving across several nodes.

Triton Inference Server

Triton Inference Server targets industrial deployment across several models.

vLLM

vLLM targets production serving, throughput and concurrency.

Unsloth

Unsloth targets fine-tuning and quantisation.

llama.cpp

llama.cpp targets workstations and modest hardware.

Method and sources

Footprints are calculated: billion parameters × bytes per parameter of the format × 1.15 runtime margin. They are not measured on hardware. The attention cache sits on top and depends on context and concurrency. Prices are indicative and not contractual. Each detail page carries its sources; the shared conventions are these.

  • Bytes per parameter: FP16 2 (16 bits per parameter, hence 2 bytes); FP8 1 (8 bits per parameter (E4M3), hence 1 byte, block scales not counted); NVFP4 0.5 (4 bits per parameter, hence 0.5 byte; 4-bit published weights weigh 0.54 to 0.56 byte per parameter with their scales (DeepSeek V4 Pro 865 GB for 1,599 billion, Kimi K3 1,561 GB for 2,780 billion)); Q4 0.6 (taken here as llama.cpp Q4_K_M, 0.6 byte per parameter including scales; Q4_K_S weighs 0.56, AWQ and GPTQ 0.55). Source: connaissance/faits.yaml, quantifications family, checked on 2026-08-31.
  • Runtime margin × 1.15: QDNA operating assumption, not measured: 15% above the weights for activations, buffers and fragmentation.
  • “Fits” threshold at 75% of memory: QDNA assumption, not measured, which keeps the remaining quarter for the attention cache and concurrency.