2 citations · 5 across the 7 of their papers we have counts for
1 paper · 1 filter
Mia Taylor, James Chua, Jan Betley +2
Reward hacking--where agents exploit flaws in imperfect reward functions rather than performing tasks as intended--poses risks for AI alignment. Reward hacking has been observed in…