35 citations · 36 across the 3 of their papers we have counts for
1 paper · 2 filters
Shangmin Guo, Biao Zhang, Tianlin Liu +9
Direct alignment from preferences (DAP) methods, such as DPO, have recently emerged as efficient alternatives to reinforcement learning from human feedback (RLHF), that do not requ…