4 citations · 6 across the 3 of their papers we have counts for
3 papers
Consequences of Misaligned AI
Simon Zhuang, Dylan Hadfield-Menell
AI systems often rely on two key components: a specified goal or reward function and an optimization algorithm to compute the optimal behavior for that goal. This approach is inten…
Multi-Principal Assistance Games: Definition and Collegial Mechanisms
Arnaud Fickinger, Simon Zhuang, Andrew Critch +2
We introduce the concept of a multi-principal assistance game (MPAG), and circumvent an obstacle in social choice theory, Gibbard's theorem, by using a sufficiently collegial prefe…
Multi-Principal Assistance Games
Arnaud Fickinger, Simon Zhuang, Dylan Hadfield-Menell +1
Assistance games (also known as cooperative inverse reinforcement learning games) have been proposed as a model for beneficial AI, wherein a robotic agent must act on behalf of a h…