Model techniques map
Techniquespost-trainingreinforcement learning algorithm

specific method · filed under post-training

Progressive Scaling of RL Data, Tasks, and Rollouts

Progressively increasing the data, tasks, and rollouts used in RL to extend capabilities across domains.

source
1
model
1
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

progressively scale the data, tasks, and rollouts employed during RL, thereby extending the model's capabilities across textual, multimodal, and agentic domains

usedpost trainingin DeepSeek-V4.1-FlashDeepSeek

Filed alongside

Other methods under post-training :: reinforcement learning algorithm.