7 citations · 9 across the 3 of their papers we have counts for
1 paper · 1 filter
Jason Mancuso, Tomasz Kisielewski, David Lindner +1
Current reinforcement learning methods fail if the reward function is imperfect, i.e. if the agent observes reward different from what it actually receives. We study this problem w…