24 citations · 47 across the 7 of their papers we have counts for
9 papers · 1 filter
Linguistic communication as (inverse) reward design
Theodore R. Sumers, Robert D. Hawkins, Mark K. Ho +2
Natural language is an intuitive and expressive way to communicate reward information to autonomous agents. It encompasses everything from concrete instructions to abstract descrip…
Consequences of Misaligned AI
Simon Zhuang, Dylan Hadfield-Menell
AI systems often rely on two key components: a specified goal or reward function and an optimization algorithm to compute the optimal behavior for that goal. This approach is inten…
Multi-Principal Assistance Games
Arnaud Fickinger, Simon Zhuang, Dylan Hadfield-Menell +1
Assistance games (also known as cooperative inverse reinforcement learning games) have been proposed as a model for beneficial AI, wherein a robotic agent must act on behalf of a h…
Conservative Agency via Attainable Utility Preservation
Alexander Matt Turner, Dylan Hadfield-Menell, Prasad Tadepalli
Reward functions are easy to misspecify; although designers can make corrections after observing mistakes, an agent pursuing a misspecified reward function can irreversibly change…
Human-AI Learning Performance in Multi-Armed Bandits
Ravi Pandya, Sandy H. Huang, Dylan Hadfield-Menell +1
People frequently face challenging decision-making problems in which outcomes are uncertain or unknown. Artificial intelligence (AI) algorithms exist that can outperform humans at…
Legible Normativity for AI Alignment: The Value of Silly Rules
Dylan Hadfield-Menell, McKane Andrus, Gillian K. Hadfield
It has become commonplace to assert that autonomous agents will have to be built to follow human rules of behavior--social norms and laws. But human laws and norms are complex and…