H200 SXM server
HGX H200, 8 SXM GPUs

| Indicative price | about €453k excl. VAT: 8-GPU HGX H200 server with 141 GB per GPU and two Xeon Gold 6538Y+ listed at £387,785 excl. VAT by CTO Servers on 2 September 2026 (single vendor; the reseller asks for a quote to validate this price), converted at the ECB rate of £0.857 per €1 (price guide) |
|---|---|
| Memory | 141 GB HBM3e per GPU, 8 GPUs (1,128 GB) |
| Bandwidth | 4.8 TB/s per GPU |
| NVLink | 900 GB/s per GPU, full NVSwitch mesh |
| Compute | 3,958 TFLOPS FP8 and 1,979 TFLOPS FP16 per GPU with sparsity, tensor parallelism across 8 GPUs |
| Power | up to 700 W per GPU |
| Form factor | 8-GPU HGX board, 8U server |
What the machine runs
A 1128 GB pool for inference and training at scale: fine-tuning, LoRA and continued pre-training across eight GPUs meshed via NVLink. Recommended runtime: vLLM with tensor parallelism.
Positioning
The HGX H200 board makes eight GPUs work as a single accelerator via NVSwitch, for training and inference of large models. NVIDIA AI Enterprise licensing is sold per GPU.
Vendors
HGX H200 systems from Dell, HPE, Lenovo, Supermicro, ASUS and Gigabyte.
GPUDirect Storage
On platforms fitted with ConnectX cards, GPUDirect Storage technology establishes a direct path between NVMe or NVMe over Fabric storage and GPU memory, bypassing the CPU buffer. Compatible storage arrays, such as NetApp, VAST, DDN or WEKA, feed training and massive RAG at full speed.
Support
Open stack supported by QDNA and the communities. Optional NVIDIA AI Enterprise (NIM, NeMo, Triton) with a service-level commitment, per GPU.
Which models fit on H200 SXM server?
The machine offers 1,128 GB of memory. 6 of the 7 open models in the catalogue load on it, in the format shown.
| Model | Billion parameters | Most precise format that fits | Sizing page |
|---|---|---|---|
| GLM 5.2 | 744 | FP8 | GLM 5.2 on H200 SXM server |
| Kimi K2.7 Code | 1,000 | NVFP4 | Kimi K2.7 Code on H200 SXM server |
| DeepSeek V4 | 1,600 | NVFP4 | DeepSeek V4 on H200 SXM server |
| Nemotron 3 Ultra | 550 | FP8 | Nemotron 3 Ultra on H200 SXM server |
| MiniMax M3 | 428 | FP16 | MiniMax M3 on H200 SXM server |
| Qwen 3.8 27B | 27 | FP16 | Qwen 3.8 27B on H200 SXM server |
See the full sizing matrix.
Frequently asked questions
How much does a configured h200 sxm cost for an LLM?
Pricing depends on configuration and supplier. The page shows a dated public price; we provide a detailed quote after a scoping call.
Which LLM fits in a h200 sxm?
Depends on the quantization format and the input context. The table on the page indicates the recommended maximum model size.
Which runtimes support the h200 sxm?
vLLM, llama.cpp, Triton Inference Server, depending on the chosen framework. The model, format and GPU combinations measured by QDNA are published in the measurements section.