On the Sample Complexity of Stability Constrained Imitation Learning
arXiv:2102.09161
Abstract
We study the following question in the context of imitation learning for continuous control: how are the underlying stability properties of an expert policy reflected in the sample-complexity of an imitation learning task? We provide the first results showing that a surprisingly granular connection can be made between the underlying expert system's incremental gain stability, a novel measure of robust convergence between pairs of system trajectories, and the dependency on the task horizon of the resulting generalization bounds. In particular, we propose and analyze incremental gain stability constrained versions of behavior cloning and a DAgger-like algorithm, and show that the resulting sample-complexity bounds naturally reflect the underlying stability properties of the expert system. As a special case, we delineate a class of systems for which the number of trajectories needed to achieve -suboptimality is sublinear in the task horizon , and do so without requiring (strong) convexity of the loss function in the policy parameters. Finally, we conduct numerical experiments demonstrating the validity of our insights on both a simple nonlinear system for which the underlying stability properties can be easily tuned, and on a high-dimensional quadrupedal robotic simulation.
References in corpus (7)
- Frustratingly Easy Domain Adaptation
- Neural Lyapunov Control
- Deeply AggreVaTeD: Differentiable Imitation Learning for Sequential Prediction
- Learning Stability Certificates from Data
- DropoutDAgger: A Bayesian Approach to Safe Imitation Learning
- Generalization Guarantees for Imitation Learning
- Imitation Learning with Stability and Safety Guarantees
Cited by in corpus (5)
- Learning Safe Multi-Agent Control with Decentralized Neural Barrier Certificates
- Imitation Learning: Progress, Taxonomies and Challenges
- TRAIL: Near-Optimal Imitation Learning with Suboptimal Data
- Implicit Behavioral Cloning
- On Imitation Learning of Linear Control Policies: Enforcing Stability and Robustness Constraints via LMI Conditions