Dual Policy Iteration
arXiv:1805.10755
Abstract
Recently, a novel class of Approximate Policy Iteration (API) algorithms have demonstrated impressive practical performance (e.g., ExIt from [2], AlphaGo-Zero from [27]). This new family of algorithms maintains, and alternately optimizes, two policies: a fast, reactive policy (e.g., a deep neural network) deployed at test time, and a slow, non-reactive policy (e.g., Tree Search), that can plan multiple steps ahead. The reactive policy is updated under supervision from the non-reactive policy, while the non-reactive policy is improved with guidance from the reactive policy. In this work we study this Dual Policy Iteration (DPI) strategy in an alternating optimization framework and provide a convergence analysis that extends existing API theory. We also develop a special instance of this framework which reduces the update of non-reactive policies to model-based optimal control using learned local models, and provides a theoretically sound way of unifying model-free and model-based RL approaches with unknown dynamics. We demonstrate the efficacy of our approach on various continuous control Markov Decision Processes.
NeurIPS 2018; Additional related works
Cited by in corpus (11)
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- When to Trust Your Model: Model-Based Policy Optimization
- On-Policy Robot Imitation Learning from a Converging Supervisor
- On the role of planning in model-based deep reinforcement learning
- Imitation-Projected Programmatic Reinforcement Learning
- Learning Self-Correctable Policies and Value Functions from Demonstrations with Negative Sampling
- Robust Inverse Reinforcement Learning under Transition Dynamics Mismatch
- Reinforced Imitation Learning by Free Energy Principle
- Evaluating model-based planning and planner amortization for continuous control
- Environment Shaping in Reinforcement Learning using State Abstraction
- Of Moments and Matching: A Game-Theoretic Framework for Closing the Imitation Gap