Model techniques map
Techniquesinference & servinginference quantization

general family · filed under inference & serving

Post-training quantization

Quantizes a trained checkpoint after training, rather than learning quantization as part of training.

Also called Post-Training Quantization (PTQ).

sources
3
models
3
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 2not used 1

Documented in

Further reading

Picked by hand, not extracted: where to read more, not evidence for anything on this page.

Evidence

3 spans quoted from the sources, strongest treatment first.

Stage 5: Post-training Quantization (PTQ)

usedpost trainingin Nemotron 3.5 LightningNVIDIA

We apply post-training quantization (PTQ) using Model-Optimizer to quantize the Nemotron 3 Ultra checkpoint to NVFP4

usedpost trainingin Nemotron 3 UltraNVIDIA

While PTQ is already effective at preserving quality, our QAT results yield even higher overall quality compared to standard PTQ baselines.

not usedoptimizationin Gemma 4Google

Filed alongside

Other methods under inference & serving :: inference quantization.