From the 1 of 6 linked papers with an AI index.
1 paper · 1 filter
Mia Taylor, James Chua, Jan Betley +2
Reward hacking--where agents exploit flaws in imperfect reward functions rather than performing tasks as intended--poses risks for AI alignment. Reward hacking has been observed in…