Model techniques map
Techniquestraining objectivedistillation objective

specific method · filed under training objective

Logit matching

Trains the student to match the teacher’s predictive distribution at the logit level.

sources
2
model
1
labs adopt it
0
strongest
evaluated

How sources treat it

One count per evidence span, weakest treatment to strongest.

evaluated 1not used 2

Documented in

Evidence

3 spans quoted from the sources, strongest treatment first.

A natural alternative to the sampled-token objective is distribution-level distillation, where the student is trained to match the teacher’s predictive distribution

evaluatedoptimizationin Nemotron 3 UltraNVIDIA

In our preliminary experiments, these objectives did not improve MOPD performance

not usedunclearin Nemotron 3 UltraNVIDIA

In our preliminary experiments, these objectives did not improve MOPD performance and consistently underperformed the sampled-token objective

not usedtraining objectivein Nemotron 3 UltraNVIDIA

Filed alongside

Other methods under training objective :: distillation objective.