Triton Inference Server
Inference Runtime

Triton serves models from multiple frameworks on a single infrastructure. It handles dynamic batching, concurrent instances and metrics tracking. It integrates with the NVIDIA AI Enterprise stack.
Key points: Multi-framework · Dynamic batching · Concurrent instances · Metrics · NVIDIA AI Enterprise.
Role
Triton Inference Server serves models from multiple frameworks on a single infrastructure, beyond large language models alone.
How it works
It handles dynamic batching, concurrent instances and metrics tracking, with fine-grained request scheduling.
Use cases
Triton suits heterogeneous fleets that combine language models, vision models and classic models, under a single console.
Integration
It integrates with the NVIDIA AI Enterprise stack and serves models in NVFP4 format on Blackwell GPUs.