5 papers
Computationally Efficient Collaborative Communication Via Regularity-Based Coarsening
Mark Bedaywi, Scott Emmons, Nika Haghtalab +1
Our results show that the existence of a short high-utility protocol already suffices for efficient communication. In particular, in a game with possible observations and a…
Provably Optimal Learning Algorithms for Assistance Games
Nivasini Ananthakrishnan, Mark Bedaywi, Michael I. Jordan +2
This paper studies an online variant of the assistance games framework, where an informed agent and an uninformed agent repeatedly interact over timesteps to optimize a common…
Learning the Preferences of a Learning Agent
Karim Abdel Sadek, Mark Bedaywi, Rhys Gould +1
For AI systems to be useful to humans, they must understand and act in accordance with our values and preferences. Since specifying preferences is a hard task, inverse reinforcemen…
Observation Interference in Partially Observable Assistance Games
Scott Emmons, Caspar Oesterheld, Vincent Conitzer +1
We study partially observable assistance games (POAGs), a model of the human-AI value alignment problem which allows the human and the AI assistant to have partial observations. Mo…
The Partially Observable Off-Switch Game
Andrew Garber, Rohan Subramani, Linus Luu +3
A wide variety of goals could cause an AI to disable its off switch because "you can't fetch the coffee if you're dead" (Russell 2019). Prior theoretical work on this shutdown prob…