activity
20172021
most citedAI Safety Gridworlds

117 citations · 134 across the 4 of their papers we have counts for

collaborators

8 papers

cs.AI20216 cited

Causal Analysis of Agent Behavior for AI Safety

Grégoire Déletang, Jordi Grau-Moya, Miljan Martic +6

As machine learning systems become more powerful they also become increasingly unpredictable and opaque. Yet, finding human-understandable explanations of how they work is essentia…

cs.AI202010 cited

Algorithms for Causal Reasoning in Probability Trees

Tim Genewein, Tom McGrath, Grégoire Déletang +4

Probability trees are one of the simplest models of causal generative processes. They possess clean semantics and -- unlike causal Bayesian networks -- they can represent context-s…

cs.AI2020

Meta-trained agents implement Bayes-optimal agents

Vladimir Mikulik, Grégoire Delétang, Tom McGrath +4

Memory-based meta-learning is a powerful technique to build agents that adapt fast to any task within a target distribution. A previous theoretical study has argued that this remar…

cs.LG2020

Avoiding Side Effects By Considering Future Tasks

Victoria Krakovna, Laurent Orseau, Richard Ngo +2

Designing reward functions is difficult: the designer has to specify what to do (what it means to complete the task) as well as what not to do (side effects that should be avoided…

cs.LG20181 cited

Scaling shared model governance via model splitting

Miljan Martic, Jan Leike, Andrew Trask +3

Currently the only techniques for sharing governance of a deep learning model are homomorphic encryption and secure multiparty computation. Unfortunately, neither of these techniqu…

cs.LG2018

Scalable agent alignment via reward modeling: a research direction

Jan Leike, David Krueger, Tom Everitt +3

One obstacle to applying reinforcement learning algorithms to real-world problems is the lack of suitable reward functions. Designing such reward functions is difficult in part bec…