taxonomy node · level 2
preference optimization
4 methods filed at this node or below it, from the sources of 4 models.
post-training :: preference optimization
Matching aids for the classifier: RLHF; DPO; pairwise comparison; human preference alignment.
In this branch 4
Everything filed at this node or below it, with one collapsible heading per child node.
By model
Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.
| Model | techniques |
|---|---|
| NVIDIA-Nemotron-3-Ultra-550B-A55B used | Reinforcement Learning from Human Feedback used— |
| Kimi K3 used | Budget-based verbosity control used— |
| Qwen3.5-397B-A17B used | Direct Preference Optimization used— |
| gpt-oss-120b used | Deliberative alignment used— |