RAIL: A modular framework for Reinforcement-learning-based Adversarial Imitation Learning
arXiv:2105.03756
Abstract
While Adversarial Imitation Learning (AIL) algorithms have recently led to state-of-the-art results on various imitation learning benchmarks, it is unclear as to what impact various design decisions have on performance. To this end, we present here an organizing, modular framework called Reinforcement-learning-based Adversarial Imitation Learning (RAIL) that encompasses and generalizes a popular subclass of existing AIL approaches. Using the view espoused by RAIL, we create two new IfO (Imitation from Observation) algorithms, which we term SAIfO: SAC-based Adversarial Imitation from Observation and SILEM (Skeletal Feature Compensation for Imitation Learning with Embodiment Mismatch). We go into greater depth about SILEM in a separate technical report. In this paper, we focus on SAIfO, evaluating it on a suite of locomotion tasks from OpenAI Gym, and showing that it outperforms contemporaneous RAIL algorithms that perform IfO.
arXiv admin note: text overlap with arXiv:2104.07810
References in corpus (9)
- Sequence to Sequence Learning with Neural Networks
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Trust Region Policy Optimization
- Generative Adversarial Imitation Learning
- Generative Adversarial Imitation from Observation
- AlgaeDICE: Policy Gradient from Arbitrary Experience
- Imitation Learning via Off-Policy Distribution Matching
- Adversarial Soft Advantage Fitting: Imitation Learning without Policy Optimization
- Off-Policy Imitation Learning from Observations