1 paper · 1 filter
Dario Garcia-Gasulla, Adrian Tormos, Anna Arias-Duart +4
Direct Preference Optimization (DPO) is an efficient alignment technique that steers LLMs towards preferable outputs by training on preference data, bypassing the need for explicit…