Model techniques map
Techniquespost-trainingreinforcement learning algorithm

specific method · filed under post-training

Group Relative Policy Optimization

A group-relative policy optimization algorithm used for reinforcement-learning training.

Also called GRPO, GRPO (Group Relative Policy Optimization), GRPO algorithm, GRPO reinforcement learning training, Group Relative Policy Optimization (GRPO), RL with GRPO.

sources
15
models
8
labs adopt it
5
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

optional 1used 15core 1

Documented in

Further reading

Picked by hand, not extracted: where to read more, not evidence for anything on this page.

Evidence

17 spans quoted from the sources, strongest treatment first.

For DeepSeek-V3.2, we still adopt Group Relative Policy Optimization (GRPO) ... as the RL training algorithm.

coreoptimizationin DeepSeek-V3.2DeepSeek

Hy3 supports GRPO reinforcement learning training with verl

usedpost trainingin Hy3Tencent Hunyuan

Our training procedure largely follows the asynchronous GRPO algorithm with the stability optimizations proposed in NVIDIA (2026)

usedoptimizationin Nemotron 3 UltraNVIDIA

For DeepSeek-V3.2, we still adopt Group Relative Policy Optimization (GRPO) as the RL training algorithm.

usedoptimizationin DeepSeek-V3.2DeepSeek

The model underwent multi-environment reinforcement learning using asynchronous GRPO (Group Relative Policy Optimization)

usedoptimizationin Nemotron 3 UltraNVIDIA

Reinforcement Learning (RL) is applied using Group Relative Policy Optimization (GRPO)

usedpost trainingin DeepSeek-V4DeepSeek

For the RL stage, we implemented the Group Relative Policy Optimization (GRPO) algorithm

usedpost trainingin DeepSeek-V4 specialist modelsDeepSeek

Reinforcement Learning using Group Relative Policy Optimization (GRPO)

usedtraining objectivein DeepSeek-V4DeepSeek

The model underwent multi-environment reinforcement learning using GRPO (Group Relative Policy Optimization) across math, code, science, instruction following, multi-step tool use, multi-turn conversations, and structured output environments.

usedpost trainingin Nemotron 3.5 LightningNVIDIA

independent cultivation of domain-specific experts (through SFT and RL with GRPO)

usedpost trainingin DeepSeek-V4-FlashDeepSeek

asynchronous GRPO (Group Relative Policy Optimization)

usedtraining objectivein Nemotron 3 UltraNVIDIA

Our RL algorithm builds upon GRPO [40] and incorporates the IcePop technique [61] to mitigate the training-inference mismatch.

usedtraining objectivein GLM-5Z.ai

the group size in the GRPO algorithm is configured to 1

usedtraining objectivein GLM-5Z.ai

asynchronous GRPO algorithm with the stability optimizations proposed in nvidia2026nemotron3superopen

usedoptimizationin Nemotron 3 UltraNVIDIA

independent cultivation of domain-specific experts (through SFT and RL with GRPO)

usedpost trainingin DeepSeek-V4DeepSeek

GRPO

usedpost trainingin MiMo-V2-FlashXiaomi

Hy3 supports GRPO reinforcement learning training with verl

optionaloptimizationin Hy3Tencent Hunyuan

Filed alongside

Other methods under post-training :: reinforcement learning algorithm.