taxonomy node · level 2
inference quantization
64 methods filed at this node or below it, from the sources of 23 models.
inference & serving :: inference quantization
Matching aids for the classifier: FP8 quantization; NVFP4 quantization; MXFP4 weights; AWQ; GGUF quantization; W4A8; KV cache quantization; weight scaling; random Hadamard transform; QuaRot.
In this branch 64
Everything filed at this node or below it, with one collapsible heading per child node.
filed here 64
By model
Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.