Model techniques map
Techniquesinference & servinginference quantization

implementation detail · filed under inference & serving

Dynamic activation scaling

Scales activations dynamically in the cited FP8 quantization setup.

source
1
models
2
lab adopt it
1
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

core 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

quant_method: fp8 with e4m3 and dynamic activation scaling

coreunclearin MiMo-V2.6-Pro and MiMo-V2.6-FlashXiaomi

Filed alongside

Other methods under inference & serving :: inference quantization.