specific method · filed under post-training
Predictive Rollout Dispatch
Estimates KV demand and inference concurrency to guide admission and placement of rollouts across ranks.
- source
- 1
- models
- 2
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1core 1
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
We jointly estimate KV demand and expected inference concurrency to guide the admission and placement of new rollouts across ranks.
coreinference servingin MiMo-V2.6Xiaomi
We jointly estimate KV demand and expected inference concurrency to guide the admission and placement of new rollouts across ranks.
usedunclearin MiMo-V2.6 RL and OPD infrastructureXiaomi
Filed alongside
Other methods under post-training :: rollout & RL infrastructure.
Asynchronous reinforcement learningPartial rolloutAsynchronous RL frameworks for large-scale agent scaffolds and environment orchestrationToken-in-token-out (TITO)Asynchronous reinforcement learning infrastructureCo-located RL trainingData SchedulerDecoupled control plane and data planeDecoupling agent rollout into sandbox and worker containerDeficit-corrected schedulingLarge-scale asynchronous RL in synthesized tasksOne-step off-policy asynchronous reinforcement learningSample-grained garbage collectionSeamless Rollout EngineSLIMEToken-granularity persistence of rollout statesToken-level interruptionTool ManagerToolboxAdaptive Rollout ConcurrencyAdaptive Rollout SchedulingAgent LoopAgent-centric rollout executionAsynchronous Agent RL algorithms