Model techniques map
Taxonomypost-trainingagentic post-training

taxonomy node · level 2

agentic post-training

31 methods filed at this node or below it, from the sources of 14 models.

post-training :: agentic post-training

Matching aids for the classifier: agentic reinforcement learning; multi-environment RLVR; tool-integrated reasoning training; multi-turn collaboration simulator.

In this branch 31

Everything filed at this node or below it, with one collapsible heading per child node.

filed here 31

Agentic reinforcement learning core · 4 sources · 4 quotes
Multi-environment reinforcement learning core · 1 source · 1 quote
Reinforcement learning in an agent harness core · 1 source · 1 quote
Agentic tool-use training used · 2 sources · 2 quotes
Environment hardening used · 2 sources · 2 quotes
Agentic post-training used · 1 source · 1 quote
Autonomous Execution Tasks (AET) used · 1 source · 1 quote
Concurrent multi-environment post-training used · 1 source · 1 quote
Container-level network isolation used · 1 source · 1 quote
Generate-verify-refine loop used · 1 source · 1 quote
Harness-optimized training used · 1 source · 1 quote
Multi-harness training used · 1 source · 3 quotes
Multi-turn collaboration simulator used · 1 source · 1 quote
Python tool use in chain-of-thought used · 1 source · 1 quote
Re-post-training for agentic capabilities used · 1 source · 1 quote
Repair-agent environment correction loop used · 1 source · 2 quotes
Seed-based terminal task generation used · 1 source · 1 quote

By model

Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.

Modeltechniques
DeepSeek-V4.1-Flash usedRepair-agent environment correction loop usedRepercussion signal for agent-crashed environments used—
DeepSeek-V4-Flash-0731 usedRe-post-training for agentic capabilities used—
NVIDIA-Nemotron-3-Ultra-550B-A55B usedConcurrent multi-environment post-training usedMulti-environment reinforcement learning from verifiable rewards used—
MiMo-V2.6-Flash coreAgentic reinforcement learning coreContainer-level network isolation usedEnvironment hardening usedEnvironment preparation to prevent solution leakage usedMulti-harness training usedUnified trajectory representation for agentic reinforcement learning used—
MiMo-V2.5 usedAgentic post-training usedAgentic reinforcement learning used—
GLM-5.3 usedEnvironment scaling for realistic training tasks used—
MiniMax-M3 usedMulti-turn collaboration simulator used—
DeepSeek-V3.2 usedGenerate-verify-refine loop usedReinforcement learning on synthetic agentic data usedTraining agentic policies with real-world tools usedReinforcement learning restricted to search and code environments evaluated—
Kimi K3 coreTraining on verifiable problems in agentic environments coreAutonomous Execution Tasks (AET) usedMock applications for personal-assistant reinforcement learning usedUnified white-box reinforcement-learning environment used—
Laguna-S-2.1 coreProduction-harness-matched reinforcement learning coreReinforcement learning in an agent harness coreMulti-harness trajectory training with native behavior preservation usedSeed-based terminal task generation used—
MiMo-V2.5-Pro usedAgentic post-training usedAgentic reinforcement learning used—
MiMo-V2.6-Pro coreAgentic reinforcement learning coreContainer-level network isolation usedEnvironment hardening usedEnvironment preparation to prevent solution leakage usedMulti-agent task decomposition and coordination usedMulti-harness training usedUnified trajectory representation for agentic reinforcement learning used—
NVIDIA-Nemotron-3.5-Lightning-30B-A3B coreMulti-environment reinforcement learning coreHarness-optimized training used—
gpt-oss-120b usedAgentic tool-use training usedPython tool use in chain-of-thought used—