specific method · filed under optimization
BF16 mixed-precision training
Training uses BF16 mixed precision throughout, with FP32 master weights.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- default
How sources treat it
One count per evidence span, weakest treatment to strongest.
default 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
We use BF16 mixed precision training throughout all stages with the master weights in FP32
defaultoptimizationin LAGUNA M.1 and LAGUNA XS.2Poolside
Filed alongside
Other methods under optimization :: training precision.
NVFP4BF16NVFP4 pre-trainingFP8 mixed-precision trainingFP4+FP8 mixed precisionFP8-precision reinforcement learningMXFP8BF16 gradient reductionBlock-wise FP8 activation quantization with offloadE2M1FP32 attention-output retentionFP32 gradient reductionFP8 storage for the residual stateHigh-precision final network layersMixed-FP8 quantizationMixed-precision trainingNVFP4 fine-grained micro-block scalingTwo-dimensional block quantization