1 paper
Zhiyu An, Duaa Nakshbandi, Wan Du
Reinforcement learning from human feedback (RLHF) implicitly aggregates heterogeneous human preferences into a single utility function, even though the underlying utilities of the…