Model techniques map
Techniquesoptimizationtraining precision

implementation detail · filed under optimization

NVFP4

NVFP4 is a 4-bit training or quantization format used for weights, activations, or gradients.

Also called NVFP4 4-bit Training Format, ultraefficient 4-bit NVFP4 training format, NVFP4 precision format, NVFP4 quantization, NVFP4 checkpoint, NVFP4 quantized checkpoints.

sources
10
models
4
labs adopt it
3
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

optional 3used 7default 1core 2

Documented in

Further reading

Picked by hand, not extracted: where to read more, not evidence for anything on this page.

Evidence

13 spans quoted from the sources, strongest treatment first.

The pre-training phase used an NVFP4 recipe.

corepost trainingin Nemotron 3.5 LightningNVIDIA

The majority of linear layers use NVFP4 for weights, activations, and gradients

coremodel architecturein Nemotron 3 UltraNVIDIA

These figures use the final NVFP4 weights that NVIDIA recommends for inference

defaultinference servingin Nemotron 3 UltraNVIDIA

use NVIDIA’s ultraefficient 4-bit NVFP4 training format on the NVIDIA Blackwell architecture

usedoptimizationin NVIDIA Nemotron 3 Super and UltraNVIDIA

NVFP4 quantized checkpoints

usedinference servingin Nemotron 3 UltraNVIDIA

We release a single NVFP4 checkpoint for Nemotron 3 Ultra

usedpost trainingin Nemotron 3 UltraNVIDIA

Pretrained in NVFP4.

usedotherin Nemotron 3 UltraNVIDIA

open-sourcing the training recipes, data, and RL environments. Checkpoints ... NVFP4 quantized

usedsoftware implementationin Nemotron 3 UltraNVIDIA

We apply post-training quantization (PTQ) using Model-Optimizer to quantize the Nemotron 3 Ultra checkpoint to NVFP4

usedpost trainingin Nemotron 3 UltraNVIDIA

We release a single NVFP4 checkpoint for Nemotron 3 Ultra

usedinference servingin Nemotron 3 UltraNVIDIA

The NVFP4 checkpoint offers a quantized alternative that reduces the aggregated VRAM requirement to at least 600 GB.

optionalinference servingin InklingThinking Machines Lab

NVIDIA publishes both BF16 and NVFP4 results across knowledge, reasoning, coding, agents, instruction following and long context.

optionalunclearin Nemotron 3.5 LightningNVIDIA

Step 3.7 Flash is also available in an NVFP4-quantized variant for efficient deployment on NVIDIA GPUs

optionalsoftware implementationin Step 3.7 FlashStepFun

Filed alongside

Other methods under optimization :: training precision.