general family · not yet filed
Scalable reinforcement learning across million-agent environments
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
core 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
Reinforcement learning scaled across million-agent environments with progressively complex task distributions for robust real-world adaptability.
corepost trainingin Qwen3.5-122B-A10BQwen