Model techniques map
Techniquesinference & servinginference quantization

implementation detail · filed under inference & serving

MXFP8 block-scale regrouping at load

Regroups block scales into MXFP8 when loading on Ascend 950 while leaving weights at one byte per element.

source
1
model
1
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

On 950 the block scales are re-grouped to MXFP8 at load (weights stay 1 byte/element).

usedinference servingin GLM-5.3-FlashZ.ai

Filed alongside

Other methods under inference & serving :: inference quantization.