activity
20162023
most citedOn the Utility of Learning about Humans for Human-AI Coordination

91 citations · 242 across the 32 of their papers we have counts for

collaborators
Showing 2021Show all

16 papers · 1 filter

cs.LG20213 cited

B-Pref: Benchmarking Preference-Based Reinforcement Learning

Kimin Lee, Laura Smith, Anca Dragan +1

Reinforcement learning (RL) requires access to a reward function that incentivizes the right behavior, but these are notoriously hard to specify for complex tasks. Preference-based…

cs.RO20211 cited

Safety Assurances for Human-Robot Interaction via Confidence-aware Game-theoretic Human Models

Ran Tian, Liting Sun, Andrea Bajcsy +2

An outstanding challenge with safety methods for human-robot interaction is reducing their conservatism while maintaining robustness to variations in human behavior. In this work,…

cs.CV20214 cited

Pragmatic Image Compression for Human-in-the-Loop Decision-Making

Siddharth Reddy, Anca D. Dragan, Sergey Levine

Standard lossy image compression algorithms aim to preserve an image's appearance, while minimizing the number of bits needed to transmit it. However, the amount of information act…

cs.RO2021

Physical Interaction as Communication: Learning Robot Objectives Online from Human Corrections

Dylan P. Losey, Andrea Bajcsy, Marcia K. O'Malley +1

When a robot performs a task next to a human, physical interaction is inevitable: the human might push, pull, twist, or guide the robot. The state-of-the-art treats these interacti…

cs.LG20214 cited

The MineRL BASALT Competition on Learning from Human Feedback

Rohin Shah, Cody Wild, Steven H. Wang +10

The last decade has seen a significant increase of interest in deep learning research, with many public successes that have demonstrated its potential. As such, these systems are n…

cs.LG20211 cited

Policy Gradient Bayesian Robust Optimization for Imitation Learning

Zaynah Javed, Daniel S. Brown, Satvik Sharma +5

The difficulty in specifying rewards for many real-world problems has led to an increased focus on learning rewards from human feedback, such as demonstrations. However, there are…