specific method · filed under post-training
Decoupled control plane and data plane
Separates trajectory payload handling from scheduling control, allowing large trajectory payloads to be buffered and transferred independently.
Also called Disaggregated Data Plane and Control Plane.
- source
- 1
- models
- 2
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
decouple the control plane from the data plane to buffer and transfer massive trajectories carrying routing and multi-modal payloads
usedsoftware implementationin MiMo-V2.6 RL infrastructureXiaomi
We therefore disaggregate the data plane from the control plane, splitting each sequence at rollout finish
usedunclearin MiMo-V2.6 RL and OPD infrastructureXiaomi
Filed alongside
Other methods under post-training :: rollout & RL infrastructure.
Asynchronous reinforcement learningPartial rolloutAsynchronous RL frameworks for large-scale agent scaffolds and environment orchestrationToken-in-token-out (TITO)Asynchronous reinforcement learning infrastructureCo-located RL trainingData SchedulerDecoupling agent rollout into sandbox and worker containerDeficit-corrected schedulingLarge-scale asynchronous RL in synthesized tasksOne-step off-policy asynchronous reinforcement learningPredictive Rollout DispatchSample-grained garbage collectionSeamless Rollout EngineSLIMEToken-granularity persistence of rollout statesToken-level interruptionTool ManagerToolboxAdaptive Rollout ConcurrencyAdaptive Rollout SchedulingAgent LoopAgent-centric rollout executionAsynchronous Agent RL algorithms