2 citations · 2 across the 3 of their papers we have counts for
1 paper
Dhawal Gupta, Adam Fisch, Christoph Dann +1
This work tackles the problem of overoptimization in reinforcement learning from human feedback (RLHF), a prevalent technique for aligning models with human preferences. RLHF relie…