Zeroth-Order Supervised Policy Improvement
arXiv:2006.06600
Abstract
Policy gradient (PG) algorithms have been widely used in reinforcement learning (RL). However, PG algorithms rely on exploiting the value function being learned with the first-order update locally, which results in limited sample efficiency. In this work, we propose an alternative method called Zeroth-Order Supervised Policy Improvement (ZOSPI). ZOSPI exploits the estimated value function globally while preserving the local exploitation of the PG methods based on zeroth-order policy optimization. This learning paradigm follows Q-learning but overcomes the difficulty of efficiently operating argmax in continuous action space. It finds max-valued action within a small number of samples. The policy learning of ZOSPI has two steps: First, it samples actions and evaluates those actions with a learned value estimator, and then it learns to perform the action with the highest value through supervised learning. We further demonstrate such a supervised learning framework can learn multi-modal policies. Experiments show that ZOSPI achieves competitive results on the continuous control benchmarks with a remarkable sample efficiency.
References in corpus (12)
- Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction
- Q-Prop: Sample-Efficient Policy Gradient with An Off-Policy Critic
- Relative Entropy Regularized Policy Iteration
- Gradientless Descent: High-Dimensional Zeroth-Order Optimization
- V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control
- Better Exploration with Optimistic Actor-Critic
- Learning to Reach Goals via Iterated Supervised Learning
- Q-Learning for Continuous Actions with Cross-Entropy Guided Policies
- Policy Continuation with Hindsight Inverse Dynamics
- Efficient Model-Free Reinforcement Learning Using Gaussian Process
- Generalized Off-Policy Actor-Critic
- Policy Search by Target Distribution Learning for Continuous Control