7 citations · 13 across the 8 of their papers we have counts for
1 paper · 2 filters
Jiacheng Liang, Yao Ma, Tharindu Kumarage +5
Reinforcement Learning from Human Feedback (RLHF) is central to aligning Large Language Models (LLMs), yet it introduces a critical vulnerability: an imperfect Reward Model (RM) ca…