117 citations · 134 across the 4 of their papers we have counts for
5 papers · 1 filter
Avoiding Side Effects By Considering Future Tasks
Victoria Krakovna, Laurent Orseau, Richard Ngo +2
Designing reward functions is difficult: the designer has to specify what to do (what it means to complete the task) as well as what not to do (side effects that should be avoided…
Scaling shared model governance via model splitting
Miljan Martic, Jan Leike, Andrew Trask +3
Currently the only techniques for sharing governance of a deep learning model are homomorphic encryption and secure multiparty computation. Unfortunately, neither of these techniqu…
Scalable agent alignment via reward modeling: a research direction
Jan Leike, David Krueger, Tom Everitt +3
One obstacle to applying reinforcement learning algorithms to real-world problems is the lack of suitable reward functions. Designing such reward functions is difficult in part bec…
Penalizing side effects using stepwise relative reachability
Victoria Krakovna, Laurent Orseau, Ramana Kumar +2
How can we design safe reinforcement learning agents that avoid unnecessary disruptions to their environment? We show that current approaches to penalizing side effects can introdu…
AI Safety Gridworlds
Jan Leike, Miljan Martic, Victoria Krakovna +5
We present a suite of reinforcement learning environments illustrating various safety properties of intelligent agents. These problems include safe interruptibility, avoiding side…