Model techniques map
Techniquesinference & servinginference quantization

specific method · filed under inference & serving

Block-scaled INT8 quantization with stochastic rounding

Uses block-scaled INT8 quantization with stochastic rounding for cache storage.

Also called Block-scaled INT8 quantization.

sources
2
model
1
labs adopt it
0
strongest
evaluated

How sources treat it

One count per evidence span, weakest treatment to strongest.

evaluated 2

Documented in

Evidence

2 spans quoted from the sources, strongest treatment first.

Block-scaled INT8 quantization with stochastic rounding largely preserves FP32-cache accuracy

evaluatedpost trainingin Nemotron 3 SuperNVIDIA

Block-scaled INT8 quantization with stochastic rounding largely preserves FP32-cache accuracy

evaluatedinference servingin Nemotron 3 SuperNVIDIA

Filed alongside

Other methods under inference & serving :: inference quantization.