specific method · filed under post-training
Prompt-level dispatch
Dispatches a new prompt when one GRPO group finishes, a strategy reported to stall on long-tail samples within a group.
Also called prompt-level dispatch of rollout prompts.
- source
- 1
- model
- 1
- labs adopt it
- 0
- strongest
- not used
How sources treat it
One count per evidence span, weakest treatment to strongest.
not used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
we switched to prompt-level dispatch, where a new prompt is dispatched after one GRPO group finishes, but found that it stalled easily on long-tail samples within a GRPO group
not usedpost trainingin DeepSeek-V4.1DeepSeek
Filed alongside
Other methods under post-training :: rollout & RL infrastructure.
Asynchronous reinforcement learningPartial rolloutAsynchronous RL frameworks for large-scale agent scaffolds and environment orchestrationToken-in-token-out (TITO)Asynchronous reinforcement learning infrastructureCo-located RL trainingData SchedulerDecoupled control plane and data planeDecoupling agent rollout into sandbox and worker containerDeficit-corrected schedulingLarge-scale asynchronous RL in synthesized tasksOne-step off-policy asynchronous reinforcement learningPredictive Rollout DispatchSample-grained garbage collectionSeamless Rollout EngineSLIMEToken-granularity persistence of rollout statesToken-level interruptionTool ManagerToolboxAdaptive Rollout ConcurrencyAdaptive Rollout SchedulingAgent LoopAgent-centric rollout execution