Model techniques map
Techniquespost-trainingreinforcement learning algorithm

specific method · filed under post-training

GRPO with Masked Importance Sampling

GRPO augmented with masked importance sampling to account for discrepancies between training and rollout policies.

source
1
model
1
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

we use GRPO (Shao et al., 2024) with masked importance sampling to account for discrepancies between the training and rollout policies

usedpost trainingin Nemotron 3 familyNVIDIA

Filed alongside

Other methods under post-training :: reinforcement learning algorithm.