Model techniques map
Techniquesoptimizationtraining precision

implementation detail · filed under optimization

MXFP8

The MXFP8 format is used for selected layers, including Mamba output projections, to preserve information or support training stability.

Also called keep these layers in MXFP8.

sources
2
models
2
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 2

Documented in

Evidence

2 spans quoted from the sources, strongest treatment first.

select layers (including latent projections, MTP layers, QKV/attention projections, and embeddings) are maintained in BF16 or MXFP8 for training stability

usedmodel architecturein Nemotron 3 UltraNVIDIA

To prevent loss of information, we keep these layers in MXFP8.

usedoptimizationin Nemotron 3 familyNVIDIA

Filed alongside

Other methods under optimization :: training precision.