specific method · filed under post-training
IcePop
A token-level masking technique incorporated into an RL algorithm to mitigate training-inference mismatch.
Also called IcePop technique, token-level masking using IcePop strategy.
- sources
- 2
- models
- 2
- labs adopt it
- 2
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 3
Documented in
Evidence
3 spans quoted from the sources, strongest treatment first.
Our RL algorithm builds upon GRPO [40] and incorporates the IcePop technique [61] to mitigate the training-inference mismatch.
usedtraining objectivein GLM-5Z.ai
where represents token-level masking using IcePop strategy (ring1t)
usedoptimizationin Nemotron 3 UltraNVIDIA
icePop strategy (ring1t)
usedoptimizationin Nemotron 3 UltraNVIDIA
Filed alongside
Other methods under post-training :: reinforcement learning algorithm.
Group Relative Policy OptimizationReinforcement LearningGroupwise Advantage RedistributionFreezing the MoE Router During Reinforcement LearningOff-Policy Sequence MaskingReinforcement Learning from Verifiable RewardsRollout Routing ReplayUnbiased KL EstimateAsynchronous Group Relative Policy OptimizationChain-of-Thought Reinforcement LearningDirect Double-Sided Importance SamplingKeep Sampling MaskMixed Reinforcement LearningPivot Reinforcement LearningReinforcement Learning Post-TrainingAbstention TrainingAdvantage ShapingAgentic RL Task MixCISPO with Length-Weighted Leave-One-Out Group-Relative AdvantagesConcatenated Routing ReplayDerived-Latency PenaltyDomain-Specialized RL ExpertsDomain-Specific GRPO TrainingDropping All-Zero-Advantage Groups