1 citations · 1 across the 1 of their papers we have counts for
1 paper
Lichang Chen, Chen Zhu, Davit Soselia +6
In this work, we study the issue of reward hacking on the response length, a challenge emerging in Reinforcement Learning from Human Feedback (RLHF) on LLMs. A well-formatted, verb…