1 paper
Jilong Liu, Yonghui Yang, Pengyang Shao +5
Direct Preference Optimization (DPO) has become a standard framework for safety alignment, but its reliance on pairwise preference updates makes training sensitive to imperfect sup…