QDNAAdvisory and architecture for LLM inference and training platforms, on-premises or hybrid

How much memory does Nemotron 3 Ultra need in Q4?

Nemotron 3 Ultra has 550 billion parameters. The Q4 format decides how much memory is needed to load them.

Short answer. In Q4, the weights of Nemotron 3 Ultra take about 379 GB including the runtime margin. 7 of the 8 platforms in the catalogue have enough memory.

How is this footprint calculated?

Nemotron 3 Ultra totals 550 billion parameters. The Q4 format takes 0.6 byte per parameter (taken here as llama.cpp Q4_K_M, 0.6 byte per parameter including scales; Q4_K_S weighs 0.56, AWQ and GPTQ 0.55). The product gives the weights, to which a 15% runtime margin is added for activations and buffers: 550 billion parameters × 0.6 byte (Q4) = 330.0 GB of weights; × 1.15 runtime margin = 379.5 GB. The attention cache is not calculated on this page. These models’ architectures (compressed latent, linear or sparse attention, Mamba layers) share no per-token formula; it is measured on the machine, with the target context and concurrency [TO BE MEASURED].

What does the format change?

Q4 brings block quantisation, broadly supported.

FormatWeights in memory Compatible platforms
FP161,265 GB3
FP8632 GB6
NVFP4316 GB7
Q4379 GB7

Which platforms qualify?

See the Nemotron 3 Ultra fact sheet.

How much memory for Nemotron 3 Ultra in Q4?

About 379 GB for the weights, excluding the attention cache.

Which format should I choose for Nemotron 3 Ultra?

Q4 brings block quantisation, broadly supported. The most precise format that fits the target machine remains the best choice.

Method and sources

Memory footprints are calculated, not measured: a reading on real hardware may differ depending on the engine and the exact weight format. Every figure below carries its source and the date of its last check.