Model techniques map
Techniquesinference & servinginference quantization

implementation detail · filed under inference & serving

Mixed-precision INT4/INT8 layer-wise quantization

Assigns INT4 to the first 30 layers and INT8 to the final 10 layers in the cited setup.

Also called mixed-precision quantization strategy.

source
1
model
1
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

We adopted a mixed-precision quantization strategy: the first 30 layers were quantized to INT4, while the final 10 layers were quantized to INT8 with 1 × 128 group-based weight gl in-smoking.

usedpost trainingin Laguna XS.2Poolside

Filed alongside

Other methods under inference & serving :: inference quantization.