Reinforcement and Imitation Learning via Interactive No-Regret Learning
arXiv:1406.5979
Abstract
Recent work has demonstrated that problems-- particularly imitation learning and structured prediction-- where a learner's predictions influence the input-distribution it is tested on can be naturally addressed by an interactive approach and analyzed using no-regret online learning. These approaches to imitation learning, however, neither require nor benefit from information about the cost of actions. We extend existing results in two directions: first, we develop an interactive imitation learning approach that leverages cost information; second, we extend the technique to address reinforcement learning. The results provide theoretical support to the commonly observed successes of online approximate policy iteration. Our approach suggests a broad new family of algorithms and provides a unifying view of existing techniques for imitation and reinforcement learning.
14 pages. Under review for NIPS 2014 conference
References in corpus (3)
Cited by in corpus (33)
- Learning to Search Better Than Your Teacher
- Deeply AggreVaTeD: Differentiable Imitation Learning for Sequential Prediction
- Learning Algorithms for Active Learning
- Global overview of Imitation Learning
- Provably Efficient Imitation Learning from Observation Alone
- Learning Belief Representations for Imitation Learning in POMDPs
- Batch Policy Learning under Constraints
- Learning to Search for Dependencies
- Learning Heuristic Search via Imitation
- Meta-Learning for Contextual Bandit Exploration
- Riemannian Motion Policy Fusion through Learnable Lyapunov Function Reshaping
- Convergence of Value Aggregation for Imitation Learning
- Uncertainty-sensitive Learning and Planning with Ensembles
- Provable Representation Learning for Imitation Learning via Bi-level Optimization
- Imitation Learning with Recurrent Neural Networks
- Learning by Cheating
- Co-training for Policy Learning
- Recruitment-imitation Mechanism for Evolutionary Reinforcement Learning
- RMP2: A Structured Composable Policy Class for Robot Learning
- Diluted Near-Optimal Expert Demonstrations for Guiding Dialogue Stochastic Policy Optimisation
- Adaptive Information Gathering via Imitation Learning
- Comparing Human-Centric and Robot-Centric Sampling for Robot Deep Learning from Demonstrations
- Smooth Imitation Learning via Smooth Costs and Smooth Policies
- Learning to Gather Information via Imitation
- Fixing exposure bias with imitation learning needs powerful oracles
- Reparameterized Variational Divergence Minimization for Stable Imitation
- Active Imitation Learning from Multiple Non-Deterministic Teachers: Formulation, Challenges, and Algorithms
- Hierarchical Variational Imitation Learning of Control Programs
- Imitation Learning via Simultaneous Optimization of Policies and Auxiliary Trajectories
- EDITOR: an Edit-Based Transformer with Repositioning for Neural Machine Translation with Soft Lexical Constraints
- Multi-Preference Actor Critic
- Trajectory-based Learning for Ball-in-Maze Games
- Concurrent Training Improves the Performance of Behavioral Cloning from Observation