implementation detail · filed under inference & serving
NVFP4 KV-cache quantization
Stores KV-cache data in NVFP4 rather than a higher-precision format.
Also called NVFP4 KV cache, NVFP4 KV caching.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- default
How sources treat it
One count per evidence span, weakest treatment to strongest.
optional 1default 1
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
All deployment snippets below default to port 8000, with chunked prefill, NVFP4 KV caching, and MTP (5 speculative tokens) enabled
defaultinference servingin Nemotron 3 UltraNVIDIA
--kv-cache-dtype nvfp4
optionalinference servingin Nemotron 3 UltraNVIDIA
Filed alongside
Other methods under inference & serving :: inference quantization.
FP8 KV-cache quantizationFP8FP4 KV-cache quantizationNVFP4 quantizationFour-Over-SixMXFP4 weight quantizationPost-training quantizationQuantizationBlock-scaled INT8 quantization with stochastic roundingFP8 E4M3 quantizationGGUFMax-based scalingMSE-based scalingMXFP8 activation quantizationNVFP4 ModelOpt re-quantizationSSM cache quantizationAWQ INT4 weight quantization (W4A16)BF16 inferenceBlock-wise E4M3 FP8 weight quantizationChannel-wise quantizationDynamic activation scalingEmbedding and KV-cache quantizationEmpirical bits-per-element budget selectionFine-grained FP8 quantization