taxonomy node · level 2
reinforcement learning algorithm
62 methods filed at this node or below it, from the sources of 18 models.
post-training :: reinforcement learning algorithm
Matching aids for the classifier: GRPO; group relative policy optimization; MLAPO; IcePop; pivot RL; direct double-sided importance sampling; rollout routing replay.
In this branch 62
Everything filed at this node or below it, with one collapsible heading per child node.
filed here 62
By model
Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.