Home › Glossary › DPO

DPO

Direct preference optimization, a fine-tuning technique that trains on pairs of better and worse responses without needing a reward model. Records contain input, preferred_output and non_preferred_output, and the two outputs can't be identical.

Also called direct preference optimization.

Read more: Microsoft Learn

In the Ultra Transcenders books

AI-300

Each book explains DPO in context, with comparison tables and the common traps.

Terms in this definition

Related terms

See DPO in the full glossary