How much memory does GLM 5.2 need in NVFP4?
GLM 5.2 has 744 billion parameters. The NVFP4 format decides how much memory is needed to load them.
How is this footprint calculated?
GLM 5.2 totals 744 billion parameters. The NVFP4 format takes 0.5 byte per parameter. The product gives the weights, to which a 15% runtime margin is added for activations and buffers. The attention cache is not included: it grows with context and concurrency.
What does the format change?
NVFP4 brings Blackwell format, native FP4 compute.
| Format | Weights in memory | Compatible platforms |
|---|---|---|
| FP16 | 1,711 GB | 2 |
| FP8 | 856 GB | 4 |
| NVFP4 | 428 GB | 7 |
| Q4 | 428 GB | 7 |
Which platforms qualify?
- Mac Studio Ultra, 512 GB
- DGX Station, 748 GB
- RTX PRO 6000 server, 768 GB
- H200 SXM server, 1,128 GB
- B200 SXM, 1,440 GB
- B300 SXM, 2,304 GB
See the GLM 5.2 fact sheet.
How much memory for GLM 5.2 in NVFP4?
About 428 GB for the weights, excluding the attention cache.
Which format should I choose for GLM 5.2?
NVFP4 brings Blackwell format, native FP4 compute. The most precise format that fits the target machine remains the best choice.
Method
Parameter counts and memory figures come from the site fact sheets. Memory footprints are calculated, not measured: a reading on real hardware may differ depending on the engine and the exact weight format. Prices are indicative and not contractual.