How much memory does Kimi K2.7 Code need in Q4?
Kimi K2.7 Code has 1000 billion parameters. The Q4 format decides how much memory is needed to load them.
How is this footprint calculated?
Kimi K2.7 Code totals 1000 billion parameters. The Q4 format takes 0.5 byte per parameter. The product gives the weights, to which a 15% runtime margin is added for activations and buffers. The attention cache is not included: it grows with context and concurrency.
What does the format change?
Q4 brings block quantisation, broadly supported.
| Format | Weights in memory | Compatible platforms |
|---|---|---|
| FP16 | 2,300 GB | 2 |
| FP8 | 1,150 GB | 3 |
| NVFP4 | 575 GB | 6 |
| Q4 | 575 GB | 6 |
Which platforms qualify?
- DGX Station, 748 GB
- RTX PRO 6000 server, 768 GB
- H200 SXM server, 1,128 GB
- B200 SXM, 1,440 GB
- B300 SXM, 2,304 GB
- GB300 NVL72, 20,700 GB
See the Kimi K2.7 Code fact sheet.
How much memory for Kimi K2.7 Code in Q4?
About 575 GB for the weights, excluding the attention cache.
Which format should I choose for Kimi K2.7 Code?
Q4 brings block quantisation, broadly supported. The most precise format that fits the target machine remains the best choice.
Method
Parameter counts and memory figures come from the site fact sheets. Memory footprints are calculated, not measured: a reading on real hardware may differ depending on the engine and the exact weight format. Prices are indicative and not contractual.