117 citations · 183 across the 10 of their papers we have counts for
4 papers · 1 filter
Avoiding Tampering Incentives in Deep RL via Decoupled Approval
Jonathan Uesato, Ramana Kumar, Victoria Krakovna +3
How can we design agents that pursue a given objective when all feedback mechanisms are influenceable by the agent? Standard RL algorithms assume a secure reward function, and can…
REALab: An Embedded Perspective on Tampering
Ramana Kumar, Jonathan Uesato, Richard Ngo +3
This paper describes REALab, a platform for embedded agency research in reinforcement learning (RL). REALab is designed to model the structure of tampering problems that may arise…
Scalable agent alignment via reward modeling: a research direction
Jan Leike, David Krueger, Tom Everitt +3
One obstacle to applying reinforcement learning algorithms to real-world problems is the lack of suitable reward functions. Designing such reward functions is difficult in part bec…
AI Safety Gridworlds
Jan Leike, Miljan Martic, Victoria Krakovna +5
We present a suite of reinforcement learning environments illustrating various safety properties of intelligent agents. These problems include safe interruptibility, avoiding side…