Model techniques map
Taxonomyoptimizationtraining precision

taxonomy node · level 2

training precision

19 methods filed at this node or below it, from the sources of 12 models.

optimization :: training precision

Matching aids for the classifier: FP8 mixed precision training; NVFP4 pre-training; MXFP8; BF16; FP32 gradient reduction; FP4+FP8 mixed precision.

In this branch 19

Everything filed at this node or below it, with one collapsible heading per child node.

filed here 19

NVFP4 core · 10 sources · 13 quotes
NVFP4 pre-training core · 5 sources · 5 quotes
BF16 mixed-precision training default · 1 source · 1 quote
BF16 used · 5 sources · 5 quotes
FP8 mixed-precision training used · 4 sources · 4 quotes
FP4+FP8 mixed precision used · 2 sources · 2 quotes
FP8-precision reinforcement learning used · 2 sources · 2 quotes
MXFP8 used · 2 sources · 2 quotes
E2M1 used · 1 source · 1 quote
FP32 attention-output retention used · 1 source · 1 quote
FP32 gradient reduction used · 1 source · 1 quote
FP8 storage for the residual state used · 1 source · 1 quote
High-precision final network layers used · 1 source · 1 quote
Mixed-FP8 quantization used · 1 source · 1 quote
Mixed-precision training used · 1 source · 1 quote
NVFP4 fine-grained micro-block scaling used · 1 source · 1 quote
Two-dimensional block quantization used · 1 source · 1 quote
BF16 gradient reduction not used · 1 source · 1 quote

By model

Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.