Maximum Entropy Deep Inverse Reinforcement Learning
arXiv:1507.04888
Abstract
This paper presents a general framework for exploiting the representational capacity of neural networks to approximate complex, nonlinear reward functions in the context of solving the inverse reinforcement learning (IRL) problem. We show in this context that the Maximum Entropy paradigm for IRL lends itself naturally to the efficient training of deep architectures. At test time, the approach leads to a computational complexity independent of the number of demonstrations, which makes it especially well-suited for applications in life-long learning scenarios. Our approach achieves performance commensurate to the state-of-the-art on existing benchmarks while exceeding on an alternative benchmark based on highly varying reward structures. Finally, we extend the basic architecture - which is equivalent to a simplified subclass of Fully Convolutional Neural Networks (FCNNs) with width one - to include larger convolutions in order to eliminate dependency on precomputed spatial features and work on raw input representations.
References in corpus (3)
Cited by in corpus (77)
- A Brief Survey of Deep Reinforcement Learning
- Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
- A Survey of Deep RL and IL for Autonomous Driving Policy Learning
- Third-Person Imitation Learning
- Intelligent Inverse Treatment Planning via Deep Reinforcement Learning, a Proof-of-Principle Study in High Dose-rate Brachytherapy for Cervical Cancer
- Trajectory Forecasts in Unknown Environments Conditioned on Grid-Based Plans
- A Review of Robot Learning for Manipulation: Challenges, Representations, and Algorithms
- A deep inverse reinforcement learning approach to route choice modeling with context-dependent rewards
- Learning to Drive using Inverse Reinforcement Learning and Deep Q-Networks
- A Survey of Deep Network Solutions for Learning Control in Robotics: From Reinforcement to Imitation
- Off-road Autonomous Vehicles Traversability Analysis and Trajectory Planning Based on Deep Inverse Reinforcement Learning
- SQIL: Imitation Learning via Reinforcement Learning with Sparse Rewards
- Energy-based Legged Robots Terrain Traversability Modeling via Deep Inverse Reinforcement Learning
- Improving Efficiency of Training a Virtual Treatment Planner Network via Knowledge-guided Deep Reinforcement Learning for Intelligent Automatic Treatment Planning of Radiotherapy
- Online Observer-Based Inverse Reinforcement Learning
- Better-than-Demonstrator Imitation Learning via Automatically-Ranked Demonstrations
- Extrapolating Beyond Suboptimal Demonstrations via Inverse Reinforcement Learning from Observations
- Where Do You Think You're Going?: Inferring Beliefs about Dynamics from Behavior
- SA-Net: Deep Neural Network for Robot Trajectory Recognition from RGB-D Streams
- AI-Driven Day-to-Day Route Choice
- Path Planning using Neural A* Search
- Inverse reinforcement learning for video games
- Robot eye-hand coordination learning by watching human demonstrations: a task function approximation approach
- Deep Reinforcement Learning and Transportation Research: A Comprehensive Review
- Internal Model from Observations for Reward Shaping
- DSDNet: Deep Structured self-Driving Network
- Fully Convolutional Search Heuristic Learning for Rapid Path Planners
- Incorporating Human Domain Knowledge into Large Scale Cost Function Learning
- A Framework and Method for Online Inverse Reinforcement Learning
- Markov Decision Process for MOOC users behavioral inference
- Learning a Prior over Intent via Meta-Inverse Reinforcement Learning
- Learning Deep Mean Field Games for Modeling Large Population Behavior
- Learning Reward Models for Cooperative Trajectory Planning with Inverse Reinforcement Learning and Monte Carlo Tree Search
- Lifelong Robotic Reinforcement Learning by Retaining Experiences
- Learning Social Navigation from Demonstrations with Conditional Neural Processes
- PixelRL: Fully Convolutional Network with Reinforcement Learning for Image Processing
- Meta Inverse Reinforcement Learning via Maximum Reward Sharing for Human Motion Analysis
- Model-based Behavioral Cloning with Future Image Similarity Learning
- A Function Approximation Method for Model-based High-Dimensional Inverse Reinforcement Learning
- Autonomous Assessment of Demonstration Sufficiency via Bayesian Inverse Reinforcement Learning
- Density Matching Reward Learning
- PsiPhi-Learning: Reinforcement Learning with Demonstrations using Successor Features and Inverse Temporal Difference Learning
- Triple-GAIL: A Multi-Modal Imitation Learning Framework with Generative Adversarial Nets
- Neural Policy Style Transfer
- f-IRL: Inverse Reinforcement Learning via State Marginal Matching
- End-to-end Interpretable Neural Motion Planner
- Replacing Rewards with Examples: Example-Based Policy Search via Recursive Classification
- Deep Inverse Q-learning with Constraints
- Human-in-the-Loop Methods for Data-Driven and Reinforcement Learning Systems
- Inverse Reinforcement Learning in Large State Spaces via Function Approximation
- Inverse Reinforcement Learning with Conditional Choice Probabilities
- Domain-Robust Visual Imitation Learning with Mutual Information Constraints
- Automatic Data Augmentation by Learning the Deterministic Policy
- Curriculum Design for Teaching via Demonstrations: Theory and Applications
- Language Conditioned Imitation Learning over Unstructured Data
- Robust Inverse Reinforcement Learning under Transition Dynamics Mismatch
- Inferring agent objectives at different scales of a complex adaptive system
- Dueling RL: Reinforcement Learning with Trajectory Preferences
- Generalized Maximum Causal Entropy for Inverse Reinforcement Learning
- Learn to Exceed: Stereo Inverse Reinforcement Learning with Concurrent Policy Optimization
- Fitting a Linear Control Policy to Demonstrations with a Kalman Constraint
- PAC-Bayesian Soft Actor-Critic Learning
- Maximum Entropy Multi-Task Inverse RL
- Inverse Reinforcement Learning via Matching of Optimality Profiles
- Predicting Goal-directed Human Attention Using Inverse Reinforcement Learning
- Learning to Optimize via Wasserstein Deep Inverse Optimal Control
- Does Unpredictability Influence Driving Behavior?
- Balancing Performance and Human Autonomy with Implicit Guidance Agent
- Sample Efficient Social Navigation Using Inverse Reinforcement Learning
- Objective-aware Traffic Simulation via Inverse Reinforcement Learning
- Deep PQR: Solving Inverse Reinforcement Learning using Anchor Actions
- Cost Functions for Robot Motion Style
- Learning rewards for robotic ultrasound scanning using probabilistic temporal ranking
- Sequential Anomaly Detection using Inverse Reinforcement Learning
- Learning Data-Driven Objectives to Optimize Interactive Systems
- Integration of Imitation Learning using GAIL and Reinforcement Learning using Task-achievement Rewards via Probabilistic Graphical Model
- Stochastic Inverse Reinforcement Learning