1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Zihui Zhao, Zechang Li
Direct Preference Optimization (DPO) has emerged as a lightweight and effective alternative to Reinforcement Learning from Human Feedback (RLHF) and Reinforcement Learning with AI…