1 paper
Zimu Lu, Aojun Zhou, Ke Wang +5
Direct Preference Optimization (DPO) has proven effective at improving the performance of large language models (LLMs) on downstream tasks such as reasoning and alignment. In this…