7 citations · 16 across the 6 of their papers we have counts for
1 paper · 1 filter
André Barreto, Vincent Dumoulin, Yiran Mao +6
Reinforcement learning from human feedback usually models preferences using a reward function that does not distinguish between people. We argue that this is unlikely to be a good…