1 paper
Nathaniel Haynam, Adam Khoja, Dhruv Kumar +2
When reward functions are hand-designed, deep reinforcement learning algorithms often suffer from reward misspecification, causing them to learn suboptimal policies in terms of the…