117 citations · 134 across the 4 of their papers we have counts for
8 papers
Causal Analysis of Agent Behavior for AI Safety
Grégoire Déletang, Jordi Grau-Moya, Miljan Martic +6
As machine learning systems become more powerful they also become increasingly unpredictable and opaque. Yet, finding human-understandable explanations of how they work is essentia…
Algorithms for Causal Reasoning in Probability Trees
Tim Genewein, Tom McGrath, Grégoire Déletang +4
Probability trees are one of the simplest models of causal generative processes. They possess clean semantics and -- unlike causal Bayesian networks -- they can represent context-s…
Meta-trained agents implement Bayes-optimal agents
Vladimir Mikulik, Grégoire Delétang, Tom McGrath +4
Memory-based meta-learning is a powerful technique to build agents that adapt fast to any task within a target distribution. A previous theoretical study has argued that this remar…
Avoiding Side Effects By Considering Future Tasks
Victoria Krakovna, Laurent Orseau, Richard Ngo +2
Designing reward functions is difficult: the designer has to specify what to do (what it means to complete the task) as well as what not to do (side effects that should be avoided…
Scaling shared model governance via model splitting
Miljan Martic, Jan Leike, Andrew Trask +3
Currently the only techniques for sharing governance of a deep learning model are homomorphic encryption and secure multiparty computation. Unfortunately, neither of these techniqu…
Scalable agent alignment via reward modeling: a research direction
Jan Leike, David Krueger, Tom Everitt +3
One obstacle to applying reinforcement learning algorithms to real-world problems is the lack of suitable reward functions. Designing such reward functions is difficult in part bec…