Model techniques map
Taxonomypost-trainingrollout & RL infrastructure

taxonomy node · level 2

rollout & RL infrastructure

76 methods filed at this node or below it, from the sources of 13 models.

post-training :: rollout & RL infrastructure

Matching aids for the classifier: asynchronous RL infrastructure; partial rollout; seamless rollout engine; SLIME; multi-task rollout orchestrator; preemptible rollout service.

In this branch 76

Everything filed at this node or below it, with one collapsible heading per child node.

filed here 76

Asynchronous reinforcement learning core · 6 sources · 7 quotes
Agent Loop core · 1 source · 1 quote
Agent-centric rollout execution core · 1 source · 1 quote
Harness Pool core · 1 source · 1 quote
Hierarchical trajectory data organization core · 1 source · 1 quote
Large-scale asynchronous RL core · 1 source · 1 quote
Payload Porter core · 1 source · 1 quote
Predictive Rollout Dispatch core · 1 source · 2 quotes
Sample-level dispatch default · 1 source · 1 quote
Partial rollout used · 3 sources · 5 quotes
Co-located RL training used · 2 sources · 2 quotes
Seamless Rollout Engine used · 2 sources · 2 quotes
SLIME used · 2 sources · 2 quotes
Token-in-token-out (TITO) used · 2 sources · 3 quotes
Adaptive Rollout Concurrency used · 1 source · 1 quote
Adaptive Rollout Scheduling used · 1 source · 1 quote
Asynchronous Agent RL algorithms used · 1 source · 1 quote
Asynchronous sample generation used · 1 source · 1 quote
Bounding the maximum off-policy ratio used · 1 source · 1 quote
Capping trajectory staleness used · 1 source · 1 quote
Closed-loop multi-turn rollout used · 1 source · 1 quote
Co-located long-context agentic RL system used · 1 source · 1 quote
Composite early-stop strategy used · 1 source · 1 quote
Compute-node sharding into scale units used · 1 source · 1 quote
Consistent configuration transitions used · 1 source · 1 quote
Data re-sampling strategy used · 1 source · 1 quote
Data Scheduler used · 1 source · 2 quotes
Decoupled control plane and data plane used · 1 source · 2 quotes
Dialogue prefix matching used · 1 source · 1 quote
Discarding early-returned short samples used · 1 source · 1 quote
Dynamic training recipe reconfiguration used · 1 source · 1 quote
Full-lifecycle monitoring of RL tasks used · 1 source · 1 quote
Heterogeneous Agent Harnesses used · 1 source · 1 quote
Incremental image transfer used · 1 source · 1 quote
Incremental multimodal-delta transfer used · 1 source · 1 quote
Isolated resettable rollout sandboxes used · 1 source · 1 quote
Long-horizon rollout budgets used · 1 source · 1 quote
Mixed-task rollout infrastructure used · 1 source · 1 quote
Multi-harness rollouts used · 1 source · 1 quote
Multi-Task Rollout Orchestrator used · 1 source · 1 quote
Multimodal rollout data handling used · 1 source · 1 quote
On-policy rollouts used · 1 source · 1 quote
Per-dataset rollout concurrency limiting used · 1 source · 1 quote
Per-source oversampling allocation used · 1 source · 1 quote
Persistent host-actor pools used · 1 source · 1 quote
Preemptible rollout service used · 1 source · 1 quote
Sample Mixer used · 1 source · 1 quote
Sample replay used · 1 source · 1 quote
Sample-grained garbage collection used · 1 source · 2 quotes
Scaffold-agnostic rollout control layer used · 1 source · 1 quote
Token-level interruption used · 1 source · 2 quotes
Tool Manager used · 1 source · 2 quotes
Toolbox used · 1 source · 2 quotes
Asynchronous group-wise grading optional · 1 source · 1 quote
Deficit-corrected scheduling evaluated · 1 source · 2 quotes
Batch-level rollout dispatch not used · 1 source · 1 quote
Prompt-level dispatch not used · 1 source · 1 quote

By model

Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.

Modeltechniques
DeepSeek-V4.1-Flash coreAsynchronous reinforcement learning framework coreSample-level dispatch defaultAsynchronous sample generation usedBounding the maximum off-policy ratio usedCo-located RL training usedCompute-node sharding into scale units usedConsistent configuration transitions usedDecoupling agent rollout into sandbox and worker container usedDiscarding early-returned short samples usedDynamic training recipe reconfiguration usedFull-lifecycle monitoring of RL tasks usedIncremental image transfer usedLarge-scale asynchronous RL in synthesized tasks usedPer-dataset rollout concurrency limiting usedPer-task upper bound on in-flight rollout samples usedPreemption-safe suspension and offloading of rollout execution usedSample-grained garbage collection usedScaffold-agnostic rollout control layer usedScaling RL along training compute and number of scaffolds usedToken-granularity persistence of rollout states usedToken-level interruption usedBatch-level rollout dispatch not usedPrompt-level dispatch not used—
NVIDIA-Nemotron-3-Ultra-550B-A55B coreOne-step off-policy asynchronous reinforcement learning coreAsynchronous reinforcement learning usedOn-policy rollouts used—
MiMo-V2.6-Flash coreAgent Loop coreAgent-centric rollout execution coreAsynchronous reinforcement learning coreHarness Pool coreHierarchical trajectory data organization corePayload Porter corePredictive Rollout Dispatch coreAdaptive Rollout Concurrency usedAdaptive Rollout Scheduling usedComposite early-stop strategy usedDecoupled control plane and data plane usedDialogue prefix matching usedHeterogeneous Agent Harnesses usedIncremental multimodal-delta transfer usedIsolated resettable rollout sandboxes usedMixed-task rollout infrastructure usedMultimodal rollout data handling usedOffloading blocking environment and tokenization work to background threads usedPartial rollout usedPer-source oversampling allocation usedPer-source priors for rollout sequence-length estimation usedPersistent host-actor pools usedSample Mixer usedSample replay usedTraining–inference consistency mechanisms usedAsynchronous group-wise grading optionalDeficit-corrected scheduling evaluated—
MiMo-V2.5 usedContinuous rollout with asynchronous reward computation and early termination usedData re-sampling strategy usedData Scheduler usedPartial rollout usedSeamless Rollout Engine usedStaleness-aware truncated importance sampling usedTool Manager usedToolbox used—
GLM-5.2 usedAsynchronous Agent RL algorithms usedAsynchronous reinforcement learning usedAsynchronous reinforcement learning infrastructure usedMulti-Task Rollout Orchestrator usedSLIME usedToken-in-token-out (TITO) used—
DeepSeek-V4-Pro usedPreemptible rollout service used—
Inkling coreLarge-scale asynchronous RL core—
Kimi K3 usedCo-located long-context agentic RL system usedCo-located RL training usedPartial rollout used—
Laguna-S-2.1 usedBlocking in-flight rollout steps on weight updates usedCapping trajectory staleness usedClosed-loop multi-turn rollout usedExact rollout-to-production chat-template alignment assertion usedLong-horizon rollout budgets usedMulti-harness rollouts usedRL sandboxing with selective network blocking and artifact caching usedToken-in-token-out (TITO) used—
MiMo-V2.5-Pro usedContinuous rollout with asynchronous reward computation and early termination usedData re-sampling strategy usedSeamless Rollout Engine used—
MiMo-V2.6-Pro coreAgent Loop coreAgent-centric rollout execution coreAsynchronous reinforcement learning coreHarness Pool coreHierarchical trajectory data organization corePayload Porter corePredictive Rollout Dispatch coreAdaptive Rollout Concurrency usedAdaptive Rollout Scheduling usedComposite early-stop strategy usedDecoupled control plane and data plane usedDialogue prefix matching usedHeterogeneous Agent Harnesses usedIncremental multimodal-delta transfer usedIsolated resettable rollout sandboxes usedMixed-task rollout infrastructure usedMixing tasks and harnesses in a single RL batch usedMultimodal rollout data handling usedOffloading blocking environment and tokenization work to background threads usedPartial rollout usedPer-source oversampling allocation usedPer-source priors for rollout sequence-length estimation usedPersistent host-actor pools usedSample Mixer usedSample replay usedTraining–inference consistency mechanisms usedAsynchronous group-wise grading optionalDeficit-corrected scheduling evaluated—
NVIDIA-Nemotron-3.5-Lightning-30B-A3B usedAsynchronous reinforcement learning used—
Qwen3.5-397B-A17B usedAsynchronous RL frameworks for large-scale agent scaffolds and environment orchestration used—