How much memory does Nemotron 3 Ultra need in Q4?
Nemotron 3 Ultra has 550 billion parameters. The Q4 format decides how much memory is needed to load them.
How is this footprint calculated?
Nemotron 3 Ultra totals 550 billion parameters. The Q4 format takes 0.5 byte per parameter. The product gives the weights, to which a 15% runtime margin is added for activations and buffers. The attention cache is not included: it grows with context and concurrency.
What does the format change?
Q4 brings block quantisation, broadly supported.
| Format | Weights in memory | Compatible platforms |
|---|---|---|
| FP16 | 1,265 GB | 3 |
| FP8 | 632 GB | 6 |
| NVFP4 | 316 GB | 7 |
| Q4 | 316 GB | 7 |
Which platforms qualify?
- Mac Studio Ultra, 512 GB
- DGX Station, 748 GB
- RTX PRO 6000 server, 768 GB
- H200 SXM server, 1,128 GB
- B200 SXM, 1,440 GB
- B300 SXM, 2,304 GB
See the Nemotron 3 Ultra fact sheet.
How much memory for Nemotron 3 Ultra in Q4?
About 316 GB for the weights, excluding the attention cache.
Which format should I choose for Nemotron 3 Ultra?
Q4 brings block quantisation, broadly supported. The most precise format that fits the target machine remains the best choice.
Method
Parameter counts and memory figures come from the site fact sheets. Memory footprints are calculated, not measured: a reading on real hardware may differ depending on the engine and the exact weight format. Prices are indicative and not contractual.