117 citations · 183 across the 10 of their papers we have counts for
11 papers · 1 filter
Alignment of Language Agents
Zachary Kenton, Tom Everitt, Laura Weidinger +3
For artificial intelligence to be beneficial to humans the behaviour of AI agents needs to be aligned with what humans want. In this paper we discuss some behavioural issues for la…
Agent Incentives: A Causal Perspective
Tom Everitt, Ryan Carey, Eric Langlois +2
We present a framework for analysing agent incentives using causal influence diagrams. We establish that a well-known criterion for value of information is complete. We propose a n…
How RL Agents Behave When Their Actions Are Modified
Eric D. Langlois, Tom Everitt
Reinforcement learning in complex environments may require supervision to prevent the agent from attempting dangerous actions. As a result of supervisor intervention, the executed…
Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective
Tom Everitt, Marcus Hutter, Ramana Kumar +1
Can humans get arbitrarily capable reinforcement learning (RL) agents to do their bidding? Or will sufficiently capable RL agents always find ways to bypass their intended objectiv…
Modeling AGI Safety Frameworks with Causal Influence Diagrams
Tom Everitt, Ramana Kumar, Victoria Krakovna +1
Proposals for safe AGI systems are typically made at the level of frameworks, specifying how the components of the proposed system should be trained and interact with each other. I…
AGI Safety Literature Review
Tom Everitt, Gary Lea, Marcus Hutter
The development of Artificial General Intelligence (AGI) promises to be a major event. Along with its many potential benefits, it also raises serious safety concerns (Bostrom, 2014…