1 paper
Yuchen Zhu, Daniel Augusto de Souza, Zhengyan Shi +4
We address the problem of reward hacking, where maximising a proxy reward does not necessarily increase the true reward. This is a key concern for Large Language Models (LLMs), as…