RTX PRO 6000 server
Air-cooled GPU server, Blackwell Server Edition

| Indicative price | on quotation: no manufacturer publishes a price for an 8-card server; only the card itself is listed, $16,000 excl. VAT on the NVIDIA marketplace as of 1 September 2026 according to thundercompute.com (price guide) |
|---|---|
| Memory | 96 GB GDDR7 per GPU (2 to 8 GPUs, up to 768 GB) |
| Compute | 4 PFLOPS FP4 with sparsity per GPU (2 PFLOPS FP8) |
| Bandwidth | 1,597 GB/s per card, PCIe Gen5 |
| Isolation | MIG up to four 24 GB instances |
| Power draw | up to 600 W per GPU (configurable), air-cooled |
| Density | 2 to 8 GPUs per server |
What the machine runs
Several models from 70 to 200 billion parameters in parallel, RAG and multi-GPU fine-tuning. Recommended runtime: vLLM. Backbone of the dwarfstar stack.
Positioning
The sovereign workhorse runs air-cooled on standard servers, with the best cost-per-token ratio.
Manufacturers
Certified servers from Dell, HPE, Lenovo, Supermicro, ASUS and Gigabyte.
GPUDirect Storage
On platforms fitted with ConnectX cards, GPUDirect Storage technology establishes a direct path between NVMe or NVMe over Fabric storage and GPU memory, bypassing the CPU buffer. Compatible arrays, such as NetApp, VAST, DDN or WEKA, feed training and large-scale RAG at full speed.
Support
Open stack supported by QDNA and the communities. Optional NVIDIA AI Enterprise (NIM, NeMo, Triton) with a service-level agreement, licensed per GPU.
Which models fit on RTX PRO 6000 server?
The machine offers 768 GB of memory. 5 of the 7 open models in the catalogue load on it, in the format shown.
| Model | Billion parameters | Most precise format that fits | Sizing page |
|---|---|---|---|
| GLM 5.2 | 744 | NVFP4 | GLM 5.2 on RTX PRO 6000 server |
| Kimi K2.7 Code | 1,000 | NVFP4 | Kimi K2.7 Code on RTX PRO 6000 server |
| Nemotron 3 Ultra | 550 | FP8 | Nemotron 3 Ultra on RTX PRO 6000 server |
| MiniMax M3 | 428 | FP8 | MiniMax M3 on RTX PRO 6000 server |
| Qwen 3.8 27B | 27 | FP16 | Qwen 3.8 27B on RTX PRO 6000 server |
See the full sizing matrix.
Frequently asked questions
How much does a configured rtx pro 6000 cost for an LLM?
Pricing depends on configuration and supplier. No public price is listed for this configuration; we provide a detailed quote after a scoping call.
Which LLM fits in a rtx pro 6000?
Depends on the quantization format and the input context. The table on the page indicates the recommended maximum model size.
Which runtimes support the rtx pro 6000?
vLLM, llama.cpp, Triton Inference Server, depending on the chosen framework. The model, format and GPU combinations measured by QDNA are published in the measurements section.