specific method · filed under post-training
Asynchronous RL frameworks for large-scale agent scaffolds and environment orchestration
Asynchronous RL infrastructure supporting large-scale agent scaffolds and environment orchestration.
- sources
- 3
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 3
Documented in
Evidence
3 spans quoted from the sources, strongest treatment first.
asynchronous RL frameworks supporting massive-scale agent scaffolds and environment orchestration
usedsoftware implementationin Qwen3.5-122B-A10BQwen
asynchronous RL frameworks supporting massive-scale agent scaffolds and environment orchestration
usedpost trainingin Qwen3.5-397B-A17BQwen
asynchronous RL frameworks supporting massive-scale agent scaffolds and environment orchestration.
usedpost trainingin Qwen3.5-35B-A3BQwen
Filed alongside
Other methods under post-training :: rollout & RL infrastructure.
Asynchronous reinforcement learningPartial rolloutToken-in-token-out (TITO)Asynchronous reinforcement learning infrastructureCo-located RL trainingData SchedulerDecoupled control plane and data planeDecoupling agent rollout into sandbox and worker containerDeficit-corrected schedulingLarge-scale asynchronous RL in synthesized tasksOne-step off-policy asynchronous reinforcement learningPredictive Rollout DispatchSample-grained garbage collectionSeamless Rollout EngineSLIMEToken-granularity persistence of rollout statesToken-level interruptionTool ManagerToolboxAdaptive Rollout ConcurrencyAdaptive Rollout SchedulingAgent LoopAgent-centric rollout executionAsynchronous Agent RL algorithms