14 citations · 14 across the 1 of their papers we have counts for
1 paper
Mohammad Gheshlaghi Azar, Mark Rowland, Bilal Piot +4
The prevalent deployment of learning from human preferences through reinforcement learning (RLHF) relies on two important approximations: the first assumes that pairwise preference…