specific method · filed under optimization
FP4 Quantization
Quantization of model weights or components to the MXFP4 format; the evidence does not establish that this is QAT.
Also called FP4 (MXFP4) quantization, MXFP4 quantization, MXFP4 quantization of MoE weights.
- sources
- 4
- models
- 3
- labs adopt it
- 2
- strongest
- default
How sources treat it
One count per evidence span, weakest treatment to strongest.
Documented in
Further reading
Picked by hand, not extracted: where to read more, not evidence for anything on this page.
- Microscaling Data Formats for Deep Learning (Rouhani et al., 2023) paper arxiv.org
- OCP Microscaling Formats (MX) Specification v1.0 docs www.opencompute.org
Evidence
4 spans quoted from the sources, strongest treatment first.
The models were post-trained with MXFP4 quantization of the MoE weights
We post-trained the models with quantization of the MoE weights to MXFP4 format
MXFP4 quantization: The models were post-trained with MXFP4 quantization of the MoE weights, making gpt-oss-120b run on a single 80GB GPU
We apply FP4 (MXFP4) quantization [OCP_MXFormat] to two components
Filed alongside
Other methods under optimization :: quantization-aware training.