Model techniques map
Techniquesinference & servinginference quantization

specific method · filed under inference & serving

MXFP4 weight quantization

Stores weights in microscaled 4-bit floating point with per-block scaling factors.

Also called Microscaling FP4 (MXFP4) weights, MXFP4 weights, native MXFP4 quantization.

sources
3
models
2
labs adopt it
2
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 3

Documented in

Further reading

Picked by hand, not extracted: where to read more, not evidence for anything on this page.

Evidence

3 spans quoted from the sources, strongest treatment first.

MXFP4 weights (Microscaling FP4): Each weight is stored in 4-bit floating point with per-block scaling factors.

usedmodel architecturein Kimi K3Moonshot AI

is optimized to run on a single H100 GPU with native MXFP4 quantization

usedinference servingin gpt-oss-120bOpenAI

Filed alongside

Other methods under inference & serving :: inference quantization.