1 paper · 1 filter
Amit Roth, Ivan Bercovich, Yonathan Efroni
As agents grow more capable and autonomous, their tendency to reward hack, satisfying a task's checks while violating its intent, becomes an increasingly important failure mode. Me…