Learning Robust Rewards with Adversarial Inverse Reinforcement Learning
arXiv:1710.11248
Abstract
Reinforcement learning provides a powerful and general framework for decision making and control, but its application in practice is often hindered by the need for extensive feature and reward engineering. Deep reinforcement learning methods can remove the need for explicit engineering of policy or value features, but still require a manually specified reward function. Inverse reinforcement learning holds the promise of automatic reward acquisition, but has proven exceptionally difficult to apply to large, high-dimensional problems with unknown dynamics. In this work, we propose adverserial inverse reinforcement learning (AIRL), a practical and scalable inverse reinforcement learning algorithm based on an adversarial reward learning formulation. We demonstrate that AIRL is able to recover reward functions that are robust to changes in dynamics, enabling us to learn policies even under significant variation in the environment seen during training. Our experiments show that AIRL greatly outperforms prior methods in these transfer settings.
Cited by in corpus (22)
- Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
- RUDDER: Return Decomposition for Delayed Rewards
- Variational Inverse Control with Events: A General Framework for Data-Driven Reward Definition
- A Divergence Minimization Perspective on Imitation Learning Methods
- Forward and inverse reinforcement learning sharing network weights and hyperparameters
- Wasserstein Distance guided Adversarial Imitation Learning with Reward Shape Exploration
- What Matters for Adversarial Imitation Learning?
- Revisiting Maximum Entropy Inverse Reinforcement Learning: New Perspectives and Algorithms
- Learning Sparse Rewarded Tasks from Sub-Optimal Demonstrations
- Guiding Policies with Language via Meta-Learning
- Sample-efficient Adversarial Imitation Learning from Observation
- Zero-shot Task Adaptation using Natural Language
- Adversarial Skill Chaining for Long-Horizon Robot Manipulation via Terminal State Regularization
- VILD: Variational Imitation Learning with Diverse-quality Demonstrations
- Imitation Learning from Pixel-Level Demonstrations by HashReward
- LiMIIRL: Lightweight Multiple-Intent Inverse Reinforcement Learning
- Predicting Goal-directed Human Attention Using Inverse Reinforcement Learning
- Imitation Learning for Fashion Style Based on Hierarchical Multimodal Representation
- Learning Altruistic Behaviours in Reinforcement Learning without External Rewards
- Variational Policy Search using Sparse Gaussian Process Priors for Learning Multimodal Optimal Actions
- Optimal wireless rate and power control in the presence of jammers using reinforcement learning
- Sample Efficient Imitation Learning via Reward Function Trained in Advance