specific method · not yet filed
Scalable RL with progressively complex task distributions in million-agent environments
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
Reinforcement learning scaled across million-agent environments with progressively complex task distributions for robust real-world adaptability.
usedpost trainingin Qwen3.5-35B-A3BQwen