1 paper · 1 filter
Ãmer Veysel ÃaÄatan, Xuandong Zhao
Reward hacking, where AI systems exploit misspecified objectives to achieve high reward without satisfying intended goals, remains a central challenge in AI safety. Yet most known…