Model techniques map
Techniquesinference & servinginference quantization

specific method · filed under inference & serving

NVFP4 quantization for routed-expert GEMMs

Quantizes routed-expert GEMMs to NVFP4, with dynamic max-based activation scaling and max-calibrated 4/6 weight scaling in the cited method.

Also called NVFP4 routed-expert GEMMs.

source
1
model
1
lab adopt it
1
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

core 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

NVFP4 routed-expert GEMMs; Dynamic max-based activation scaling and max-calibrated 4/6 weight scaling.

corepost trainingin Nemotron 3 UltraNVIDIA

Filed alongside

Other methods under inference & serving :: inference quantization.