1 paper · 1 filter
Thomas Kwa, Drake Thomas, Adrià Garriga-Alonso
When applying reinforcement learning from human feedback (RLHF), the reward is learned from data and, therefore, always has some error. It is common to mitigate this by regularizin…