specific method · filed under post-training
CISPO with Length-Weighted Leave-One-Out Group-Relative Advantages
A token-level REINFORCE surrogate combining CISPO importance-sampling-ratio clipping with a length-weighted leave-one-out group-relative advantage estimator.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
core 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
We train with a token-level REINFORCE surrogate combining importance-sampling-ratio clipping (CISPO [14]) with a length-weighted leave-one-out [2] group-relative advantage estimator.
coreoptimizationin Laguna XS.2Poolside
Filed alongside
Other methods under post-training :: reinforcement learning algorithm.
Group Relative Policy OptimizationReinforcement LearningGroupwise Advantage RedistributionFreezing the MoE Router During Reinforcement LearningIcePopOff-Policy Sequence MaskingReinforcement Learning from Verifiable RewardsRollout Routing ReplayUnbiased KL EstimateAsynchronous Group Relative Policy OptimizationChain-of-Thought Reinforcement LearningDirect Double-Sided Importance SamplingKeep Sampling MaskMixed Reinforcement LearningPivot Reinforcement LearningReinforcement Learning Post-TrainingAbstention TrainingAdvantage ShapingAgentic RL Task MixConcatenated Routing ReplayDerived-Latency PenaltyDomain-Specialized RL ExpertsDomain-Specific GRPO TrainingDropping All-Zero-Advantage Groups