1 citations · 1 across the 4 of their papers we have counts for
1 paper · 1 filter
Shangmin Guo, Biao Zhang, Tianlin Liu +9
Direct alignment from preferences (DAP) methods, such as DPO, have recently emerged as efficient alternatives to reinforcement learning from human feedback (RLHF), that do not requ…