24 citations · 47 across the 7 of their papers we have counts for
19 papers
Linguistic communication as (inverse) reward design
Theodore R. Sumers, Robert D. Hawkins, Mark K. Ho +2
Natural language is an intuitive and expressive way to communicate reward information to autonomous agents. It encompasses everything from concrete instructions to abstract descrip…
What are you optimizing for? Aligning Recommender Systems with Human Values
Jonathan Stray, Ivan Vendrov, Jeremy Nixon +2
We describe cases where real recommender systems were modified in the service of various human values such as diversity, fairness, well-being, time well spent, and factual accuracy…
Consequences of Misaligned AI
Simon Zhuang, Dylan Hadfield-Menell
AI systems often rely on two key components: a specified goal or reward function and an optimization algorithm to compute the optimal behavior for that goal. This approach is inten…
Multi-Principal Assistance Games: Definition and Collegial Mechanisms
Arnaud Fickinger, Simon Zhuang, Andrew Critch +2
We introduce the concept of a multi-principal assistance game (MPAG), and circumvent an obstacle in social choice theory, Gibbard's theorem, by using a sufficiently collegial prefe…
Multi-Principal Assistance Games
Arnaud Fickinger, Simon Zhuang, Dylan Hadfield-Menell +1
Assistance games (also known as cooperative inverse reinforcement learning games) have been proposed as a model for beneficial AI, wherein a robotic agent must act on behalf of a h…
Silly rules improve the capacity of agents to learn stable enforcement and compliance behaviors
Raphael Köster, Dylan Hadfield-Menell, Gillian K. Hadfield +1
How can societies learn to enforce and comply with social norms? Here we investigate the learning dynamics and emergence of compliance and enforcement of social norms in a foraging…