Model techniques map
Techniquesinference & servinginference quantization

specific method · filed under inference & serving

Max-based scaling

Sets a quantization scale using the block’s absolute maximum.

Also called Max-based quantization scale calibration, Max-based weight scaling.

sources
2
model
1
labs adopt it
0
strongest
evaluated

How sources treat it

One count per evidence span, weakest treatment to strongest.

evaluated 2

Documented in

Evidence

2 spans quoted from the sources, strongest treatment first.

max-based scaling uses the block absolute maximum

evaluatedpost trainingin Nemotron 3 UltraNVIDIA

we experimented with max-based, MSE-based, and Four-Over-Six scaling

evaluatedpost trainingin Nemotron 3 UltraNVIDIA

Filed alongside

Other methods under inference & serving :: inference quantization.