117 citations · 211 across the 17 of their papers we have counts for
Showing 2020 · cs.LGShow all
2 papers · 2 filters
cs.LG2020★ 2 cited
Avoiding Tampering Incentives in Deep RL via Decoupled Approval
Jonathan Uesato, Ramana Kumar, Victoria Krakovna +3
How can we design agents that pursue a given objective when all feedback mechanisms are influenceable by the agent? Standard RL algorithms assume a secure reward function, and can…
cs.LG2020★ 3 cited
REALab: An Embedded Perspective on Tampering
Ramana Kumar, Jonathan Uesato, Richard Ngo +3
This paper describes REALab, a platform for embedded agency research in reinforcement learning (RL). REALab is designed to model the structure of tampering problems that may arise…