1 paper
Zae Myung Kim, Chanwoo Park, Vipul Raheja +2
Reward-based alignment methods for large language models (LLMs) face two key limitations: vulnerability to reward hacking, where models exploit flaws in the reward signal; and reli…