Reinforcement and Imitation Learning via Interactive No-Regret Learning
arXiv:1406.5979
Abstract
Recent work has demonstrated that problems-- particularly imitation learning and structured prediction-- where a learner's predictions influence the input-distribution it is tested on can be naturally addressed by an interactive approach and analyzed using no-regret online learning. These approaches to imitation learning, however, neither require nor benefit from information about the cost of actions. We extend existing results in two directions: first, we develop an interactive imitation learning approach that leverages cost information; second, we extend the technique to address reinforcement learning. The results provide theoretical support to the commonly observed successes of online approximate policy iteration. Our approach suggests a broad new family of algorithms and provides a unifying view of existing techniques for imitation and reinforcement learning.
14 pages. Under review for NIPS 2014 conference
References in corpus (3)
Cited by in corpus (10)
- Learning to Search Better Than Your Teacher
- Deeply AggreVaTeD: Differentiable Imitation Learning for Sequential Prediction
- Learning Algorithms for Active Learning
- Global overview of Imitation Learning
- Learning to Search for Dependencies
- Learning Heuristic Search via Imitation
- Convergence of Value Aggregation for Imitation Learning
- Imitation Learning with Recurrent Neural Networks
- Comparing Human-Centric and Robot-Centric Sampling for Robot Deep Learning from Demonstrations
- Learning to Gather Information via Imitation