1 paper
Hua Qu, Yifan Li, Xiaodong Yuan
Direct Preference Optimization (DPO) has become an important method for aligning large language models (LLMs) with human preferences because it removes the need for explicit reward…