Serving Nemotron 3 Ultra with vLLM
Nemotron 3 Ultra has 550 billion parameters. vLLM targets production serving, throughput and concurrency. This page gives the memory required and the platforms that qualify.
What is the minimum memory?
Nemotron 3 Ultra totals 550 billion parameters. In NVFP4 its weights take about 316 GB including the runtime margin: 550 billion parameters × 0.5 byte (NVFP4) = 275.0 GB of weights; × 1.15 runtime margin = 316.2 GB. In FP8 the footprint doubles, to about 632 GB (550 billion parameters × 1 byte (FP8) = 550.0 GB of weights; × 1.15 runtime margin = 632.5 GB.) The attention cache is not calculated on this page. These models’ architectures (compressed latent, linear or sparse attention, Mamba layers) share no per-token formula; it is measured on the machine, with the target context and concurrency [TO BE MEASURED].
What is vLLM for?
vLLM targets production serving, throughput and concurrency. See the vLLM and Nemotron 3 Ultra fact sheets.
On which platforms?
- Nemotron 3 Ultra on Mac Studio Ultra
- Nemotron 3 Ultra on DGX Station
- Nemotron 3 Ultra on RTX PRO 6000 server
- Nemotron 3 Ultra on H200 SXM server
- Nemotron 3 Ultra on B200 SXM
- Nemotron 3 Ultra on B300 SXM
How much memory for Nemotron 3 Ultra with vLLM?
About 316 GB in NVFP4 for the weights, excluding the attention cache.
Does vLLM suit Nemotron 3 Ultra?
vLLM serves Nemotron 3 Ultra as soon as the machine offers at least 316 GB of memory, the weight footprint in NVFP4. 7 of the 8 platforms in the catalogue meet that bar.
Method and sources
Memory footprints are calculated, not measured: a reading on real hardware may differ depending on the engine and the exact weight format. Every figure below carries its source and the date of its last check.
- Nemotron 3 Ultra: 550 billion parameters, 55 billion active, 262,144-token context, OpenMDW 1.1 licence. Hugging Face repository (config.json, API, model card), checked on 2 September 2026. 550 billion announced; the BF16 repository holds 560 billion elements (1,121 GB). The card claims “up to 1M” context, config.json declares 262,144.
- Bytes per parameter: NVFP4 0.5 (4 bits per parameter, hence 0.5 byte; 4-bit published weights weigh 0.54 to 0.56 byte per parameter with their scales (DeepSeek V4 Pro 865 GB for 1,599 billion, Kimi K3 1,561 GB for 2,780 billion)); FP8 1 (8 bits per parameter (E4M3), hence 1 byte, block scales not counted). Source: connaissance/faits.yaml, quantifications family, checked on 2026-08-31.
- Runtime margin × 1.15: QDNA operating assumption, not measured: 15% above the weights for activations, buffers and fragmentation.
- “Fits” threshold at 75% of memory: QDNA assumption, not measured, which keeps the remaining quarter for the attention cache and concurrency.