specific method · filed under post-training
Direct Preference Optimization
A method used to align model behavior with human preferences, distinguished here from RLHF.
Also called Direct Preference Optimization (DPO).
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
We further align model behavior with human preferences through Direct Preference Optimization (DPO) (Rafailov et al., 2023).
usedpost trainingin Qwen3.5-OmniQwen
Filed alongside
Other methods under post-training :: preference optimization.