taxonomy node · level 2
prediction head
5 methods filed at this node or below it, from the sources of 14 models.
model architecture :: prediction head
Matching aids for the classifier: multi-token prediction; MTP; MTP layer; shared-weight prediction heads; Nemotron hybrid MTP.
In this branch 5
Everything filed at this node or below it, with one collapsible heading per child node.
By model
Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.
| Model | techniques |
|---|---|
| Hy4-preview core | Multi-Token Prediction core— |
| NVIDIA-Nemotron-3-Ultra-550B-A55B core | Multi-Token Prediction coreSampled prior-MTP hidden-state conditioning usedShared-weight design across prediction heads usedNemotron Hybrid Multi-Token Prediction optional— |
| MiMo-V2.6-Flash core | Multi-Token Prediction core— |
| DeepSeek-V4-Flash core | Multi-Token Prediction core— |
| MiMo-V2.5 core | Multi-Token Prediction core— |
| Hy3 core | Multi-Token Prediction core— |
| GLM-5.2 used | Multi-Token Prediction used— |
| DeepSeek-V4-Pro core | Multi-Token Prediction core— |
| MiMo-V2.5-Pro used | Multi-Token Prediction used— |
| MiMo-V2.6-Pro core | Multi-Token Prediction core— |
| NVIDIA-Nemotron-3.5-Lightning-30B-A3B core | Multi-Token Prediction core— |
| Qwen3.5-397B-A17B used | Multi-Token Prediction usedMulti-Token Prediction for residual codebooks used— |
| Qwen3.8-Flash-Next core | Multi-Token Prediction core— |
| Step-3.7-Flash optional | Multi-Token Prediction optional— |