specific method · filed under post-training
Repair-agent environment correction loop
A repair agent fixes identified environment errors, recalibrates evaluation points, and sends the task back for verification.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
If the inspection does not pass, a repair agent fixes all identified errors, adjusts evaluation points that are too easy or too difficult, and the task re-enters verification.
usedunclearin DeepSeek-V4.1-FlashDeepSeek
a repair agent fixes all identified errors, adjusts evaluation points that are too easy or too difficult, and the task re-enters verification
useddata curationin DeepSeek-V4.1 RL training dataDeepSeek
Filed alongside
Other methods under post-training :: agentic post-training.
Agentic reinforcement learningMulti-harness trainingAgentic tool-use trainingEnvironment hardeningReinforcement learning on synthetic agentic dataAgentic post-trainingAutonomous Execution Tasks (AET)Concurrent multi-environment post-trainingContainer-level network isolationEnvironment preparation to prevent solution leakageEnvironment scaling for realistic training tasksGenerate-verify-refine loopHarness-optimized trainingMock applications for personal-assistant reinforcement learningMulti-agent task decomposition and coordinationMulti-environment reinforcement learningMulti-environment reinforcement learning from verifiable rewardsMulti-harness trajectory training with native behavior preservationMulti-turn collaboration simulatorProduction-harness-matched reinforcement learningPython tool use in chain-of-thoughtRe-post-training for agentic capabilitiesReinforcement learning in an agent harnessReinforcement learning restricted to search and code environments