specific method · filed under data curation
Knowledge distillation for synthetic data
Synthetic data is generated by distilling trajectories, solutions, or translations from strong teacher models or agent systems.
Also called knowledge distillation, distillation.
- sources
- 3
- models
- 2
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
Documented in
Evidence
3 spans quoted from the sources, strongest treatment first.
During post-training, we generate synthetic data by distilling trajectories, solutions, and translations from strong teacher models and agent systems
we generate synthetic data by distilling trajectories, solutions, and translations from strong teacher models
we generate synthetic data by distilling trajectories, solutions, and translations from strong teacher models and agent systems, often grounded in real tasks or documents and aggressively filtered for quality.
Filed alongside
Other methods under data curation :: synthetic data.