117 citations · 119 across the 3 of their papers we have counts for
7 papers
AI Safety Gridworlds
Jan Leike, Miljan Martic, Victoria Krakovna +5
We present a suite of reinforcement learning environments illustrating various safety properties of intelligent agents. These problems include safe interruptibility, avoiding side…
A Game-Theoretic Analysis of the Off-Switch Game
Tobias Wängberg, Mikael Böörs, Elliot Catt +2
The off-switch game is a game theoretic model of a highly intelligent robot interacting with a human. In the original paper by Hadfield-Menell et al. (2016), the analysis is not fu…
Count-Based Exploration in Feature Space for Reinforcement Learning
Jarryd Martin, Suraj Narayanan Sasikumar, Tom Everitt +1
We introduce a new count-based optimistic exploration algorithm for Reinforcement Learning (RL) that is feasible in environments with high-dimensional state-action spaces. The succ…
Death and Suicide in Universal Artificial Intelligence
Jarryd Martin, Tom Everitt, Marcus Hutter
Reinforcement learning (RL) is a general paradigm for studying intelligent behaviour, with applications ranging from artificial intelligence to psychology and economics. AIXI is a…
Avoiding Wireheading with Value Reinforcement Learning
Tom Everitt, Marcus Hutter
How can we design good goals for arbitrarily intelligent agents? Reinforcement learning (RL) is a natural approach. Unfortunately, RL does not work well for generally intelligent a…
Self-Modification of Policy and Utility Function in Rational Agents
Tom Everitt, Daniel Filan, Mayank Daswani +1
Any agent that is part of the environment it interacts with and has versatile actuators (such as arms and fingers), will in principle have the ability to self-modify -- for example…