general family · filed under post-training
Reinforcement Learning
Reinforcement learning is named without a more specific algorithm or mechanism in the supplied evidence.
Also called RL, Reinforcement Learning (RL), advanced reinforcement learning techniques, reinforcement learning (RL) training, strengthened reinforcement learning.
- sources
- 11
- models
- 5
- labs adopt it
- 4
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
Documented in
Further reading
Picked by hand, not extracted: where to read more, not evidence for anything on this page.
- Training language models to follow instructions with human feedback (Ouyang et al., 2022) paper arxiv.orgcanonical RLHF post-training reference
Evidence
12 spans quoted from the sources, strongest treatment first.
Through joint optimization of SFT and RL
advanced reinforcement learning techniques with concurrent multi-environment post-training at scale
scaling up RL training
post-trained using Supervised Fine Tuning (SFT), Reinforcement Learning (RL), and Multi-teacher On-Policy Distillation
Post-trained with enhanced pipeline involving Supervised Fine Tuning (SFT), Reinforcement Learning (RL), and Multi-teacher On-Policy Distillation (MOPD) for improved model accuracy.
followed by Reinforcement Learning using Group Relative Policy Optimization (GRPO)
through strengthened reinforcement learning and enhanced data quality and diversity
multi-environment reinforcement learning using asynchronous GRPO
post-trained using ... Reinforcement Learning (RL)
post-trained using SFT, RL, and MOPD
large-scale Reinforcement Learning (RL)
Filed alongside
Other methods under post-training :: reinforcement learning algorithm.