Reparameterized Variational Divergence Minimization for Stable Imitation
arXiv:2006.10810
Abstract
While recent state-of-the-art results for adversarial imitation-learning algorithms are encouraging, recent works exploring the imitation learning from observation (ILO) setting, where trajectories \textit{only} contain expert observations, have not been met with the same success. Inspired by recent investigations of -divergence manipulation for the standard imitation learning setting(Ke et al., 2019; Ghasemipour et al., 2019), we here examine the extent to which variations in the choice of probabilistic divergence may yield more performant ILO algorithms. We unfortunately find that -divergence minimization through reinforcement learning is susceptible to numerical instabilities. We contribute a reparameterization trick for adversarial imitation learning to alleviate the optimization challenges of the promising -divergence minimization framework. Empirically, we demonstrate that our design choices allow for ILO algorithms that outperform baseline approaches and more closely match expert performance in low-dimensional continuous-control tasks.
References in corpus (11)
- Energy-based Generative Adversarial Network
- Generative Moment Matching Networks
- A Connection between Generative Adversarial Networks, Inverse Reinforcement Learning, and Energy-Based Models
- Lyapunov-based Safe Policy Optimization for Continuous Control
- Reinforcement and Imitation Learning via Interactive No-Regret Learning
- Learning Invariant Feature Spaces to Transfer Skills with Reinforcement Learning
- A unified view of entropy-regularized Markov decision processes
- On integral probability metrics, ϕ-divergences and binary classification
- Deeply AggreVaTeD: Differentiable Imitation Learning for Sequential Prediction
- A Divergence Minimization Perspective on Imitation Learning Methods
- Provably Efficient Imitation Learning from Observation Alone