taxonomy node · level 2
language modelling objective
6 methods filed at this node or below it, from the sources of 10 models.
training objective :: language modelling objective
Matching aids for the classifier: next-token prediction; sampled-token objective; text pre-training.
In this branch 6
Everything filed at this node or below it, with one collapsible heading per child node.
By model
Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.
| Model | techniques |
|---|---|
| DeepSeek-V4.1-Flash used | Next-token prediction used— |
| DeepSeek-V4-Flash-0731 optional | Fill-in-the-middle (FIM) completion optional— |
| NVIDIA-Nemotron-3-Ultra-550B-A55B used | Sampled-token objective used— |
| MiMo-V2.6-Flash used | Cross-entropy objective used— |
| MiMo-V2.5 used | Text pre-training used— |
| DeepSeek-V3.2 used | Stricter token constraints during training used— |
| DeepSeek-V4-Flash-Vision-Exp optional | Fill-in-the-middle (FIM) completion optional— |
| DeepSeek-V4-Pro-0813 optional | Fill-in-the-middle (FIM) completion optional— |
| MiMo-V2.5-Pro used | Text pre-training used— |
| MiMo-V2.6-Pro used | Cross-entropy objective used— |