taxonomy node · level 2
auxiliary loss
9 methods filed at this node or below it, from the sources of 9 models.
training objective :: auxiliary loss
Matching aids for the classifier: MoE sequence auxiliary loss; KL alignment loss; gradient detach; sample-level attention masking.
In this branch 9
Everything filed at this node or below it, with one collapsible heading per child node.
By model
Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.
| Model | techniques |
|---|---|
| DeepSeek-V4.1-Flash used | Sample-level attention masking usedSequence-level balance loss used— |
| NVIDIA-Nemotron-3-Ultra-550B-A55B used | Unfinished-trajectory loss masking used— |
| MiMo-V2.6-Flash used | Loss masking used— |
| DeepSeek-V4-Flash used | Sample-level attention masking used— |
| MiMo-V2.5 used | MoE sequence auxiliary loss used— |
| MiniMax-M3 used | Gradient detachment usedKL alignment loss usedLanguage modeling loss plus KL loss used— |
| DeepSeek-V3.2 used | Domain-specific KL regularization strength usedKL alignment loss used— |
| DeepSeek-V4-Pro used | Sample-level attention masking used— |
| MiMo-V2.6-Pro used | Loss masking used— |