Model techniques map
Techniquespost-trainingreinforcement learning algorithm

specific method · filed under post-training

Keep Routing

Preserving expert routing paths used during sampling and enforcing those same paths during training.

source
1
model
1
lab adopt it
1
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

core 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

we preserve the expert routing paths used during sampling in the inference framework and enforce the same routing paths during training

coreoptimizationin DeepSeek-V3.2DeepSeek

Filed alongside

Other methods under post-training :: reinforcement learning algorithm.