Interactive Learning from Policy-Dependent Human Feedback
arXiv:1701.06049
Abstract
This paper investigates the problem of interactively learning behaviors communicated by a human teacher using positive and negative feedback. Much previous work on this problem has made the assumption that people provide feedback for decisions that is dependent on the behavior they are teaching and is independent from the learner's current policy. We present empirical results that show this assumption to be false -- whether human trainers give a positive or negative feedback for a decision is influenced by the learner's current policy. Based on this insight, we introduce {\em Convergent Actor-Critic by Humans} (COACH), an algorithm for learning from policy-dependent feedback that converges to a local optimum. Finally, we demonstrate that COACH can successfully learn multiple behaviors on a physical robot.
8 pages + references, 5 figures
Cited by in corpus (40)
- Curriculum Learning for Reinforcement Learning Domains: A Framework and Survey
- Reward Machines: Exploiting Reward Function Structure in Reinforcement Learning
- Scalable agent alignment via reward modeling: a research direction
- What Matters in Learning from Offline Human Demonstrations for Robot Manipulation
- Reward-rational (implicit) choice: A unifying formalism for reward learning
- APPLE: Adaptive Planner Parameter Learning from Evaluative Feedback
- Human-in-the-Loop Deep Reinforcement Learning with Application to Autonomous Driving
- Interactively shaping robot behaviour with unlabeled human instructions
- A Survey on Interactive Reinforcement Learning: Design Principles and Open Challenges
- The EMPATHIC Framework for Task Learning from Implicit Human Feedback
- Widening the Pipeline in Human-Guided Reinforcement Learning with Explanation and Context-Aware Data Augmentation
- Recent Advances in Leveraging Human Guidance for Sequential Decision-Making Tasks
- FRESH: Interactive Reward Shaping in High-Dimensional State Spaces using Human Feedback
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training
- Faster and Safer Training by Embedding High-Level Knowledge into Deep Reinforcement Learning
- Integrating Behavior Cloning and Reinforcement Learning for Improved Performance in Dense and Sparse Reward Environments
- Learning Human Objectives by Evaluating Hypothetical Behavior
- Neural-encoding Human Experts' Domain Knowledge to Warm Start Reinforcement Learning
- Learning Behaviors with Uncertain Human Feedback
- Skill Preferences: Learning to Extract and Execute Robotic Skills from Human Feedback
- Inverse Constrained Reinforcement Learning
- Using Machine Teaching to Investigate Human Assumptions when Teaching Reinforcement Learners
- Deep Interactive Reinforcement Learning for Path Following of Autonomous Underwater Vehicle
- Human-in-the-Loop Methods for Data-Driven and Reinforcement Learning Systems
- Human-guided Robot Behavior Learning: A GAN-assisted Preference-based Reinforcement Learning Approach
- Robot Learning via Human Adversarial Games
- X2T: Training an X-to-Text Typing Interface with Online Learning from User Feedback
- A Human-Centered Data-Driven Planner-Actor-Critic Architecture via Logic Programming
- Learning Online from Corrective Feedback: A Meta-Algorithm for Robotics
- Directed Policy Gradient for Safe Reinforcement Learning with Human Advice
- Automata Learning from Preference and Equivalence Queries
- Transfer Learning Across Simulated Robots With Different Sensors
- Teachable Reinforcement Learning via Advice Distillation
- MarioMix: Creating Aligned Playstyles for Bots with Interactive Reinforcement Learning
- A Joint Planning and Learning Framework for Human-Aided Decision-Making
- Convergence of a Human-in-the-Loop Policy-Gradient Algorithm With Eligibility Trace Under Reward, Policy, and Advantage Feedback
- Show or Tell? Demonstration is More Robust to Changes in Shared Perception than Explanation
- Combining Reward Information from Multiple Sources
- Multi-Preference Actor Critic
- Information Directed Reward Learning for Reinforcement Learning