Model techniques map
Techniquesinference & servinginference quantization

implementation detail · filed under inference & serving

NVFP4 KV-cache quantization

Stores KV-cache data in NVFP4 rather than a higher-precision format.

Also called NVFP4 KV cache, NVFP4 KV caching.

source
1
model
1
lab adopt it
1
strongest
default

How sources treat it

One count per evidence span, weakest treatment to strongest.

optional 1default 1

Documented in

Evidence

2 spans quoted from the sources, strongest treatment first.

All deployment snippets below default to port 8000, with chunked prefill, NVFP4 KV caching, and MTP (5 speculative tokens) enabled

defaultinference servingin Nemotron 3 UltraNVIDIA

--kv-cache-dtype nvfp4

optionalinference servingin Nemotron 3 UltraNVIDIA

Filed alongside

Other methods under inference & serving :: inference quantization.