specific method · filed under post-training
Deficit-corrected scheduling
Combines generation demand and remaining-batch deficit to balance collection and occupancy.
- source
- 1
- models
- 2
- labs adopt it
- 0
- strongest
- evaluated
How sources treat it
One count per evidence span, weakest treatment to strongest.
evaluated 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
Deficit-corrected scheduling (α = 0.5) improves occupancy stability over Deficit-based scheduling (α = 0) and collection balance over Target-based scheduling (α = 1) in this workload.
evaluatedunclearin MiMo-V2.6 RL and OPD infrastructure
Deficit-corrected scheduling (𝛼 = 0.5) improves occupancy stability over Deficit-based scheduling (𝛼 = 0) and collection balance over Target-based scheduling (𝛼 = 1) in this workload.
evaluatedinference servingin MiMo-V2.6Xiaomi
Filed alongside
Other methods under post-training :: rollout & RL infrastructure.
Asynchronous reinforcement learningPartial rolloutAsynchronous RL frameworks for large-scale agent scaffolds and environment orchestrationToken-in-token-out (TITO)Asynchronous reinforcement learning infrastructureCo-located RL trainingData SchedulerDecoupled control plane and data planeDecoupling agent rollout into sandbox and worker containerLarge-scale asynchronous RL in synthesized tasksOne-step off-policy asynchronous reinforcement learningPredictive Rollout DispatchSample-grained garbage collectionSeamless Rollout EngineSLIMEToken-granularity persistence of rollout statesToken-level interruptionTool ManagerToolboxAdaptive Rollout ConcurrencyAdaptive Rollout SchedulingAgent LoopAgent-centric rollout executionAsynchronous Agent RL algorithms