activity
20152017
most citedAI Safety Gridworlds

117 citations · 119 across the 3 of their papers we have counts for

collaborators

7 papers

cs.LG2017117 cited

AI Safety Gridworlds

Jan Leike, Miljan Martic, Victoria Krakovna +5

We present a suite of reinforcement learning environments illustrating various safety properties of intelligent agents. These problems include safe interruptibility, avoiding side…

cs.GT2017

A Game-Theoretic Analysis of the Off-Switch Game

Tobias Wängberg, Mikael Böörs, Elliot Catt +2

The off-switch game is a game theoretic model of a highly intelligent robot interacting with a human. In the original paper by Hadfield-Menell et al. (2016), the analysis is not fu…

cs.AI2017

Count-Based Exploration in Feature Space for Reinforcement Learning

Jarryd Martin, Suraj Narayanan Sasikumar, Tom Everitt +1

We introduce a new count-based optimistic exploration algorithm for Reinforcement Learning (RL) that is feasible in environments with high-dimensional state-action spaces. The succ…

cs.AI2016

Death and Suicide in Universal Artificial Intelligence

Jarryd Martin, Tom Everitt, Marcus Hutter

Reinforcement learning (RL) is a general paradigm for studying intelligent behaviour, with applications ranging from artificial intelligence to psychology and economics. AIXI is a…

cs.AI2016

Avoiding Wireheading with Value Reinforcement Learning

Tom Everitt, Marcus Hutter

How can we design good goals for arbitrarily intelligent agents? Reinforcement learning (RL) is a natural approach. Unfortunately, RL does not work well for generally intelligent a…

cs.AI2016

Self-Modification of Policy and Utility Function in Rational Agents

Tom Everitt, Daniel Filan, Mayank Daswani +1

Any agent that is part of the environment it interacts with and has versatile actuators (such as arms and fingers), will in principle have the ability to self-modify -- for example…