SAFARI: Safe and Active Robot Imitation Learning with Imagination
arXiv:2011.09586
Abstract
One of the main issues in Imitation Learning is the erroneous behavior of an agent when facing out-of-distribution situations, not covered by the set of demonstrations given by the expert. In this work, we tackle this problem by introducing a novel active learning and control algorithm, SAFARI. During training, it allows an agent to request further human demonstrations when these out-of-distribution situations are met. At deployment, it combines model-free acting using behavioural cloning with model-based planning to reduce state-distribution shift, using future state reconstruction as a test for state familiarity. We empirically demonstrate how this method increases the performance on a set of manipulation tasks with respect to passive Imitation Learning, by gathering more informative demonstrations and by minimizing state-distribution shift at test time. We also show how this method enables the agent to autonomously predict failure rapidly and safely.
References in corpus (7)
- Deep and Confident Prediction for Time Series at Uber
- One-Shot Visual Imitation Learning via Meta-Learning
- Trial without Error: Towards Safe Reinforcement Learning via Human Intervention
- Model-Predictive Policy Learning with Uncertainty Regularization for Driving in Dense Traffic
- DropoutDAgger: A Bayesian Approach to Safe Imitation Learning
- On the exact relationship between the denoising function and the data distribution
- Active Imitation Learning from Multiple Non-Deterministic Teachers: Formulation, Challenges, and Algorithms