Model techniques map
Techniquespost-trainingreinforcement learning algorithm

specific method · filed under post-training

Direct Double-Sided Importance Sampling

A token-level clipping or masking strategy for rollout log-probabilities that controls off-policy bias without tracking historical policy checkpoints.

source
1
model
1
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 2

Documented in

Evidence

2 spans quoted from the sources, strongest treatment first.

we employ a Direct Double-sided Importance Sampling, which applies a token-level clipping mechanism ( ) to rollout log-probabilities, while efficiently controlling off-policy bias without tracking historical policy checkpoints.

usedtraining objectivein GLM-5Z.ai

we employ a double-sided calibration token-level masking strategy

usedoptimizationin GLM-5Z.ai

Filed alongside

Other methods under post-training :: reinforcement learning algorithm.