1 paper
Shivanshu Shekhar, Shreyas Singh, Tong Zhang
Direct Preference Optimization (DPO) has been successfully used to align large language models (LLMs) according to human preferences, and more recently it has also been applied to…