4 citations · 4 across the 1 of their papers we have counts for
1 paper · 1 filter
Christopher Frye, Ilya Feige
Autonomous agents trained via reinforcement learning present numerous safety concerns: reward hacking, negative side effects, and unsafe exploration, among others. In the context o…