activity
20172022
most citedWhat are you optimizing for? Aligning Recommender Systems with Human Values

24 citations · 47 across the 7 of their papers we have counts for

collaborators
Showing cs.AIShow all

9 papers · 1 filter

cs.AI20221 cited

Linguistic communication as (inverse) reward design

Theodore R. Sumers, Robert D. Hawkins, Mark K. Ho +2

Natural language is an intuitive and expressive way to communicate reward information to autonomous agents. It encompasses everything from concrete instructions to abstract descrip…

cs.AI2021

Consequences of Misaligned AI

Simon Zhuang, Dylan Hadfield-Menell

AI systems often rely on two key components: a specified goal or reward function and an optimization algorithm to compute the optimal behavior for that goal. This approach is inten…

cs.AI20202 cited

Multi-Principal Assistance Games

Arnaud Fickinger, Simon Zhuang, Dylan Hadfield-Menell +1

Assistance games (also known as cooperative inverse reinforcement learning games) have been proposed as a model for beneficial AI, wherein a robotic agent must act on behalf of a h…

cs.AI2019

Conservative Agency via Attainable Utility Preservation

Alexander Matt Turner, Dylan Hadfield-Menell, Prasad Tadepalli

Reward functions are easy to misspecify; although designers can make corrections after observing mistakes, an agent pursuing a misspecified reward function can irreversibly change…

cs.AI2018

Human-AI Learning Performance in Multi-Armed Bandits

Ravi Pandya, Sandy H. Huang, Dylan Hadfield-Menell +1

People frequently face challenging decision-making problems in which outcomes are uncertain or unknown. Artificial intelligence (AI) algorithms exist that can outperform humans at…

cs.AI2018

Legible Normativity for AI Alignment: The Value of Silly Rules

Dylan Hadfield-Menell, McKane Andrus, Gillian K. Hadfield

It has become commonplace to assert that autonomous agents will have to be built to follow human rules of behavior--social norms and laws. But human laws and norms are complex and…