Model techniques map
Techniquesinference & servinginference quantization

general family · filed under inference & serving

Post-quantization of MoE model layers

A broad approach that quantizes MoE layers and the KV cache into formats including FP8, INT4, and NVFP4.

Also called quantization of MoE layers.

source
1
model
1
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

We quantized the model’s MoE layers to FP8, INT4, NVAP, and the KV cache to FP8.

usedpost trainingin Laguna XS.2Poolside

Filed alongside

Other methods under inference & serving :: inference quantization.