9 citations · 10 across the 3 of their papers we have counts for
1 paper · 1 filter
Zhichao Wang, Bin Bi, Can Huang +7
RL alignment methods, including RLHF and DPO, are primarily based on pairwise preference data. Although scalar or score-based feedback has been collected in some settings, it is ra…