specific method · filed under training objective
Logit matching
Trains the student to match the teacher’s predictive distribution at the logit level.
- sources
- 2
- model
- 1
- labs adopt it
- 0
- strongest
- evaluated
How sources treat it
One count per evidence span, weakest treatment to strongest.
evaluated 1not used 2
Documented in
Evidence
3 spans quoted from the sources, strongest treatment first.
A natural alternative to the sampled-token objective is distribution-level distillation, where the student is trained to match the teacher’s predictive distribution
evaluatedoptimizationin Nemotron 3 UltraNVIDIA
In our preliminary experiments, these objectives did not improve MOPD performance
not usedunclearin Nemotron 3 UltraNVIDIA
In our preliminary experiments, these objectives did not improve MOPD performance and consistently underperformed the sampled-token objective
not usedtraining objectivein Nemotron 3 UltraNVIDIA
Filed alongside
Other methods under training objective :: distillation objective.