specific method · filed under post-training
Scalable RL at Agent Scale
Reinforcement learning scaled across million-agent environments with progressively more complex task distributions.
- sources
- 3
- models
- 2
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
Documented in
Evidence
3 spans quoted from the sources, strongest treatment first.
Qwen3.5 was trained with reinforcement learning scaled across what the team describes as "million-agent environments with progressively complex task distributions."
Reinforcement learning scaled across million-agent environments with progressively complex task distributions for robust real-world adaptability.
Reinforcement learning scaled across million-agent environments with progressively complex task distributions for robust real-world adaptability.
Filed alongside
Other methods under post-training.