Model techniques map
Taxonomypost-trainingpreference optimization

taxonomy node · level 2

preference optimization

4 methods filed at this node or below it, from the sources of 4 models.

post-training :: preference optimization

Matching aids for the classifier: RLHF; DPO; pairwise comparison; human preference alignment.

In this branch 4

Everything filed at this node or below it, with one collapsible heading per child node.

filed here 4

Reinforcement Learning from Human Feedback used · 3 sources · 3 quotes
Budget-based verbosity control used · 1 source · 1 quote
Deliberative alignment used · 1 source · 1 quote
Direct Preference Optimization used · 1 source · 1 quote

By model

Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.