107 citations · 407 across the 30 of their papers we have counts for
4 papers · 1 filter
Guided Imitation of Task and Motion Planning
Michael James McDonald, Dylan Hadfield-Menell
While modern policy optimization methods can do complex manipulation from sensory data, they struggle on problems with extended time horizons and multiple sub-goals. On the other h…
Robust Feature-Level Adversaries are Interpretability Tools
Stephen Casper, Max Nadeau, Dylan Hadfield-Menell +1
The literature on adversarial attacks in computer vision typically focuses on pixel-level perturbations. These tend to be very difficult to interpret. Recent work that manipulates…
What are you optimizing for? Aligning Recommender Systems with Human Values
Jonathan Stray, Ivan Vendrov, Jeremy Nixon +2
We describe cases where real recommender systems were modified in the service of various human values such as diversity, fairness, well-being, time well spent, and factual accuracy…
Consequences of Misaligned AI
Simon Zhuang, Dylan Hadfield-Menell
AI systems often rely on two key components: a specified goal or reward function and an optimization algorithm to compute the optimal behavior for that goal. This approach is inten…