Model techniques map
Techniquesdata curationsynthetic data

specific method · filed under data curation

Knowledge distillation for synthetic data

Synthetic data is generated by distilling trajectories, solutions, or translations from strong teacher models or agent systems.

Also called knowledge distillation, distillation.

sources
3
models
2
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 3

Documented in

Evidence

3 spans quoted from the sources, strongest treatment first.

During post-training, we generate synthetic data by distilling trajectories, solutions, and translations from strong teacher models and agent systems

useddata curationin Nemotron 3.5 LightningNVIDIA

we generate synthetic data by distilling trajectories, solutions, and translations from strong teacher models

useddata curationin Nemotron 3 UltraNVIDIA

we generate synthetic data by distilling trajectories, solutions, and translations from strong teacher models and agent systems, often grounded in real tasks or documents and aggressively filtered for quality.

useddata curationin Nemotron 3.5 LightningNVIDIA

Filed alongside

Other methods under data curation :: synthetic data.