Direct preference optimization, a fine-tuning technique that trains on pairs of better and worse responses without needing a reward model. Records contain input, preferred_output and non_preferred_output, and the two outputs can't be identical.
Also called direct preference optimization.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains DPO in context, with comparison tables and the common traps.
Terms in this definition
- Fine-tuning
Continuing to train a pretrained base model on task examples, through supervised fine-tuning, DPO or RFT, so its weights shift towards a style, format or task. It neither acts as a safety control nor adds new knowledge well.
Related terms
- GPT-4.1-mini
An Azure OpenAI chat model that is quicker and cheaper, accepts images as well as text and offers a context window of roughly 1M tokens, smaller on certain deployment types. It can be fine-tuned with SFT or DPO.
- GPT-4o
OpenAI chat model that is multimodal, taking images and text and, in some variants, audio. Temperature and penalty parameters work with it, and the 2024-08-06 version allows supervised, DPO and vision fine-tuning.