Nemotron 3 Ultra
NVIDIA

Nemotron 3 Ultra is the most capable open American model. It combines 550 billion parameters, of which 55 billion are active, in a Mamba 2 architecture paired with a transformer, in NVFP4 format, with a one-million-token context. It aligns with the NVIDIA AI Enterprise stack.
Key points: Open American model · Mamba 2 plus transformer · NVFP4 · 1M context · NVIDIA AI Enterprise.
Architecture
Nemotron 3 Ultra combines 550 billion parameters, of which 55 billion are active, in a Mamba 2 architecture paired with a transformer, in NVFP4 format, with a one-million-token context.
Strengths
It stands as the reference open American model and aligns with the NVIDIA AI Enterprise stack.
Use cases
Nemotron suits long-context workloads and deployments aligned with the NVIDIA ecosystem, with enterprise support.
Deployment
Served by vLLM or Triton in NVFP4 format on Blackwell GPUs, it takes advantage of native FP4 hardware acceleration.
All models on the platform are interchangeable through a single gateway: switching models is just a matter of changing one configuration line, with no code rewrite. Local execution remains the priority.