Model techniques map
Techniquesoptimizationtraining precision

implementation detail · filed under optimization

BF16 gradient reduction

Gradient reduction or local accumulation uses BF16 precision rather than FP32.

source
1
model
1
labs adopt it
0
strongest
not used

How sources treat it

One count per evidence span, weakest treatment to strongest.

not used 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

The first divergence, which occurred at around 8T tokens, was attributed to a reduction in local gradient accumulation precision for the output layer from FP32 to BF16

not usedoptimizationin Nemotron 3 UltraNVIDIA

Filed alongside

Other methods under optimization :: training precision.