15 citations · 17 across the 7 of their papers we have counts for
1 paper · 1 filter
Xinpeng Wang, Nitish Joshi, Barbara Plank +2
Reward hacking, where a reasoning model exploits loopholes in a reward function to achieve high rewards without solving the intended task, poses a significant threat. This behavior…