Model techniques map
Techniquesinference & servinginference quantization

specific method · filed under inference & serving

FP8 E4M3 quantization

Uses the E4M3 FP8 format for quantization; the cited evidence does not specify the tensors being quantized.

sources
2
model
1
labs adopt it
0
strongest
evaluated

How sources treat it

One count per evidence span, weakest treatment to strongest.

evaluated 2

Documented in

Evidence

2 spans quoted from the sources, strongest treatment first.

FP8 E4M3 quantization degrades accuracy

evaluatedpost trainingin Nemotron 3 SuperNVIDIA

FP8 E4M3 quantization degrades accuracy

evaluatedinference servingin Nemotron 3 SuperNVIDIA

Filed alongside

Other methods under inference & serving :: inference quantization.