QDNAAdvisory and architecture for LLM inference and training platforms, on-premises or hybrid

Serving Nemotron 3 Ultra with Dynamo

Nemotron 3 Ultra has 550 billion parameters. Dynamo targets distributed serving across several nodes. This page gives the memory required and the platforms that qualify.

Short answer. Dynamo serves Nemotron 3 Ultra as soon as the machine offers at least 316 GB of memory, the weight footprint in NVFP4. 7 of the 8 platforms in the catalogue meet that bar.

What is the minimum memory?

Nemotron 3 Ultra totals 550 billion parameters. In NVFP4 its weights take about 316 GB including the runtime margin: 550 billion parameters × 0.5 byte (NVFP4) = 275.0 GB of weights; × 1.15 runtime margin = 316.2 GB. In FP8 the footprint doubles, to about 632 GB (550 billion parameters × 1 byte (FP8) = 550.0 GB of weights; × 1.15 runtime margin = 632.5 GB.) The attention cache is not calculated on this page. These models’ architectures (compressed latent, linear or sparse attention, Mamba layers) share no per-token formula; it is measured on the machine, with the target context and concurrency [TO BE MEASURED].

What is Dynamo for?

Dynamo targets distributed serving across several nodes. See the Dynamo and Nemotron 3 Ultra fact sheets.

On which platforms?

How much memory for Nemotron 3 Ultra with Dynamo?

About 316 GB in NVFP4 for the weights, excluding the attention cache.

Does Dynamo suit Nemotron 3 Ultra?

Dynamo serves Nemotron 3 Ultra as soon as the machine offers at least 316 GB of memory, the weight footprint in NVFP4. 7 of the 8 platforms in the catalogue meet that bar.

Method and sources

Memory footprints are calculated, not measured: a reading on real hardware may differ depending on the engine and the exact weight format. Every figure below carries its source and the date of its last check.