Model techniques map
Techniquesinference & servinginference quantization

ambiguous · filed under inference & serving

GGUF

The evidence identifies GGUF as a model-weight format but does not specify a quantization level or algorithm.

Also called GGUF (GPT-Generated Unified Format), GGUF quantization.

sources
2
model
1
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

optional 1used 1

Documented in

Evidence

2 spans quoted from the sources, strongest treatment first.

Released under Apache 2.0 with BF16, FP8, NVFP4, and GGUF weights on Hugging Face

usedsoftware implementationin Step 3.7 FlashStepFun

GGUF: https://huggingface.co/stepfun-ai/Step-3.7-Flash-GGUF

optionalsoftware implementationin Step 3.7 FlashStepFun

Filed alongside

Other methods under inference & serving :: inference quantization.