Model techniques map
Techniquesinference & servinginference quantization

ambiguous · filed under inference & serving

FP8

The evidence identifies FP8 as a weight format but does not specify a distinct quantization scheme.

Also called FP8 (8-bit floating point), FP8 quantization, --quantization fp8, FP8 model, FP8 quantized instruct model, Hy4 preview-FP8.

sources
14
models
9
labs adopt it
6
strongest
default

How sources treat it

One count per evidence span, weakest treatment to strongest.

optional 4used 8default 2

Documented in

Evidence

14 spans quoted from the sources, strongest treatment first.

the default zai-org/GLM-5.3 checkpoint is now native FP8

defaultinference servingin GLM-5.3Z.ai

ships native FP8 weights

defaultinference servingin GLM-5.3-FlashZ.ai

Hy3-FP8 | FP8 quantized instruct model

usedsoftware implementationin Hy3Tencent Hunyuan

Released under Apache 2.0 with BF16, FP8, NVFP4, and GGUF weights on Hugging Face

usedsoftware implementationin Step 3.7 FlashStepFun

Hy4 preview-FP8 model weights

usedpost trainingin Hy4-previewTencent Hunyuan

KV Cache is quantized to FP8

usedpost trainingin Nemotron 3 UltraNVIDIA

We recommend using the official FP8 checkpoint Qwen/Qwen3.5-397B-A17B-FP8 for optimal serving efficiency.

usedinference servingin Qwen3.5-397B-A17BQwen

Start the FP8 model

usedinference servingin Hy4-previewTencent Hunyuan

--quantization fp8

usedinference servingin MiMo-V2.5Xiaomi

For FP8 model ... vllm serve ...

usedsoftware implementationin Step 3.7 FlashStepFun

Hy4 preview-FP8 | FP8 quantized instruct model

optionalinference servingin Hy4-previewTencent Hunyuan

A separate Hy3-FP8 checkpoint is also released. FP8 lowers the memory footprint for cheaper serving.

optionalpost trainingin Hy3Tencent

Qwen's official FP8 checkpoint

optionalunclearin Qwen3.6-35B-A3BQwen

--quantization fp8

optionalinference servingin MiMo-V2.5Xiaomi

Filed alongside

Other methods under inference & serving :: inference quantization.