taxonomy node · level 2
distillation objective
7 methods filed at this node or below it, from the sources of 6 models.
training objective :: distillation objective
Matching aids for the classifier: logit matching; full-vocabulary logit distillation; reverse KL divergence; temperature-scaled forward KL.
In this branch 7
Everything filed at this node or below it, with one collapsible heading per child node.
By model
Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.
| Model | techniques |
|---|---|
| NVIDIA-Nemotron-3-Ultra-550B-A55B used | Temperature-scaled forward KL distillation usedLogit matching evaluated— |
| DeepSeek-V4-Flash used | Full-vocabulary logit distillation used— |
| MiMo-V2.5 used | Reverse KL divergence used— |
| DeepSeek-V4-Pro used | Full-vocabulary logit distillation used— |
| Laguna-S-2.1 used | Hidden-state caching for KL distillation usedQuantization-aware distillation used— |
| Qwen3.8-Flash-Next used | Dense-attention distillation used— |