53 citations · 58 across the 13 of their papers we have counts for
1 paper · 1 filter
Zhiyu An, Duaa Nakshbandi, Wan Du
Reinforcement learning from human feedback (RLHF) implicitly aggregates heterogeneous human preferences into a single utility function, even though the underlying utilities of the…