QDNASales and integration of LLM inference and training platforms, on-premises or hybrid

Nemotron 3 Ultra

NVIDIA

Nemotron 3 Ultra

Nemotron 3 Ultra is the most capable open American model. It combines 550 billion parameters, of which 55 billion are active, in a Mamba 2 architecture paired with a transformer, in NVFP4 format, with a one-million-token context. It aligns with the NVIDIA AI Enterprise stack.

Key points: Open American model · Mamba 2 plus transformer · NVFP4 · 1M context · NVIDIA AI Enterprise.

Architecture

Nemotron 3 Ultra combines 550 billion parameters, of which 55 billion are active, in a Mamba 2 architecture paired with a transformer, in NVFP4 format, with a one-million-token context.

Strengths

It stands as the reference open American model and aligns with the NVIDIA AI Enterprise stack.

Use cases

Nemotron suits long-context workloads and deployments aligned with the NVIDIA ecosystem, with enterprise support.

Deployment

Served by vLLM or Triton in NVFP4 format on Blackwell GPUs, it takes advantage of native FP4 hardware acceleration.

All models on the platform are interchangeable through a single gateway: switching models is just a matter of changing one configuration line, with no code rewrite. Local execution remains the priority.

Deploy this model on your premises

On your own hardware, with your data never leaving.

Book a call