Model techniques map
Techniquespost-training

specific method · filed under post-training

Scalable RL at Agent Scale

Reinforcement learning scaled across million-agent environments with progressively more complex task distributions.

sources
3
models
2
lab adopt it
1
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 1core 2

Documented in

Evidence

3 spans quoted from the sources, strongest treatment first.

Qwen3.5 was trained with reinforcement learning scaled across what the team describes as "million-agent environments with progressively complex task distributions."

corepost trainingin Qwen3.5-397B-A17BAlibaba

Reinforcement learning scaled across million-agent environments with progressively complex task distributions for robust real-world adaptability.

coretraining objectivein Qwen3.5Qwen

Reinforcement learning scaled across million-agent environments with progressively complex task distributions for robust real-world adaptability.

usedpost trainingin Qwen3.5-397B-A17BQwen

Filed alongside

Other methods under post-training.