Model techniques map
Techniquesinference & servinginference quantization

implementation detail · filed under inference & serving

SpinQuant R1 rotation

Applies a SpinQuant R1 rotation before quantization to improve FP8 quality without runtime overhead in the cited setup.

source
1
model
1
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

We found that using SpinQuant R1 rotation improved the quality of FP8 quantization without working any runtime overhead, and therefore adopted it for our FP8 quantization scheme.

usedpost trainingin Laguna XS.2Poolside

Filed alongside

Other methods under inference & serving :: inference quantization.