Model techniques map
Techniquespost-trainingpolicy distillation

specific method · filed under post-training

Specialist Distillation

Distillation using domain-specialized models or teachers to produce domain-specific data or a unified model.

Also called multi-domain specialist distillation, specialist-model distillation.

sources
3
models
2
labs adopt it
2
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 3

Documented in

Evidence

3 spans quoted from the sources, strongest treatment first.

For each task, we initially develop a specialized model dedicated exclusively to that particular domain

usedpost trainingin DeepSeek-V3.2DeepSeek

Once the specialist models are prepared, they are used to produce the domain-specific data for the final checkpoint.

usedpost trainingin DeepSeek-V3.2DeepSeek

we first train a suite of domain-specialized teacher models via independent Supervised Fine-Tuning (SFT) and reinforcement learning (RL).

usedpost trainingin Qwen3.5-OmniQwen

Filed alongside

Other methods under post-training :: policy distillation.