Model techniques map
Techniquespost-trainingreinforcement learning algorithm

general family · filed under post-training

Reinforcement Learning

Reinforcement learning is named without a more specific algorithm or mechanism in the supplied evidence.

Also called RL, Reinforcement Learning (RL), advanced reinforcement learning techniques, reinforcement learning (RL) training, strengthened reinforcement learning.

sources
11
models
5
labs adopt it
4
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 12

Documented in

Further reading

Picked by hand, not extracted: where to read more, not evidence for anything on this page.

Evidence

12 spans quoted from the sources, strongest treatment first.

Through joint optimization of SFT and RL

usedpost trainingin Hy3Tencent Hunyuan

advanced reinforcement learning techniques with concurrent multi-environment post-training at scale

usedpost trainingin Nemotron 3 familyNVIDIA

scaling up RL training

usedpost trainingin Hy3Tencent Hunyuan

post-trained using Supervised Fine Tuning (SFT), Reinforcement Learning (RL), and Multi-teacher On-Policy Distillation

usedpost trainingin Nemotron 3 UltraNVIDIA

Post-trained with enhanced pipeline involving Supervised Fine Tuning (SFT), Reinforcement Learning (RL), and Multi-teacher On-Policy Distillation (MOPD) for improved model accuracy.

usedpost trainingin Nemotron 3 UltraNVIDIA

followed by Reinforcement Learning using Group Relative Policy Optimization (GRPO)

usedtraining objectivein DeepSeek-V4DeepSeek

through strengthened reinforcement learning and enhanced data quality and diversity

usedtraining objectivein Hy3Tencent

multi-environment reinforcement learning using asynchronous GRPO

usedtraining objectivein Nemotron 3 UltraNVIDIA

post-trained using ... Reinforcement Learning (RL)

usedpost trainingin Nemotron 3 UltraNVIDIA

post-trained using SFT, RL, and MOPD

usedpost trainingin Nemotron 3 UltraNVIDIA

large-scale Reinforcement Learning (RL)

usedpost trainingin MiMo-V2-FlashXiaomi

RL and MOPD

usedoptimizationin MiMo-V2.5Xiaomi

Filed alongside

Other methods under post-training :: reinforcement learning algorithm.