specific method · filed under post-training
Direct Double-Sided Importance Sampling
A token-level clipping or masking strategy for rollout log-probabilities that controls off-policy bias without tracking historical policy checkpoints.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
we employ a Direct Double-sided Importance Sampling, which applies a token-level clipping mechanism ( ) to rollout log-probabilities, while efficiently controlling off-policy bias without tracking historical policy checkpoints.
usedtraining objectivein GLM-5Z.ai
we employ a double-sided calibration token-level masking strategy
usedoptimizationin GLM-5Z.ai
Filed alongside
Other methods under post-training :: reinforcement learning algorithm.
Group Relative Policy OptimizationReinforcement LearningGroupwise Advantage RedistributionFreezing the MoE Router During Reinforcement LearningIcePopOff-Policy Sequence MaskingReinforcement Learning from Verifiable RewardsRollout Routing ReplayUnbiased KL EstimateAsynchronous Group Relative Policy OptimizationChain-of-Thought Reinforcement LearningKeep Sampling MaskMixed Reinforcement LearningPivot Reinforcement LearningReinforcement Learning Post-TrainingAbstention TrainingAdvantage ShapingAgentic RL Task MixCISPO with Length-Weighted Leave-One-Out Group-Relative AdvantagesConcatenated Routing ReplayDerived-Latency PenaltyDomain-Specialized RL ExpertsDomain-Specific GRPO TrainingDropping All-Zero-Advantage Groups