117 citations · 184 across the 11 of their papers we have counts for
4 papers · 1 filter
Alignment of Language Agents
Zachary Kenton, Tom Everitt, Laura Weidinger +3
For artificial intelligence to be beneficial to humans the behaviour of AI agents needs to be aligned with what humans want. In this paper we discuss some behavioural issues for la…
Agent Incentives: A Causal Perspective
Tom Everitt, Ryan Carey, Eric Langlois +2
We present a framework for analysing agent incentives using causal influence diagrams. We establish that a well-known criterion for value of information is complete. We propose a n…
Equilibrium Refinements for Multi-Agent Influence Diagrams: Theory and Practice
Lewis Hammond, James Fox, Tom Everitt +2
Multi-agent influence diagrams (MAIDs) are a popular form of graphical model that, for certain classes of games, have been shown to offer key complexity and explainability advantag…
How RL Agents Behave When Their Actions Are Modified
Eric D. Langlois, Tom Everitt
Reinforcement learning in complex environments may require supervision to prevent the agent from attempting dangerous actions. As a result of supervisor intervention, the executed…