Combining policy gradient and Q-learning
arXiv:1611.01626
Abstract
Policy gradient is an efficient technique for improving a policy in a reinforcement learning setting. However, vanilla online variants are on-policy only and not able to take advantage of off-policy data. In this paper we describe a new technique that combines policy gradient with off-policy Q-learning, drawing experience from a replay buffer. This is motivated by making a connection between the fixed points of the regularized policy gradient algorithm and the Q-values. This connection allows us to estimate the Q-values from the action preferences of the policy, to which we apply Q-learning updates. We refer to the new technique as 'PGQL', for policy gradient and Q-learning. We also establish an equivalency between action-value fitting techniques and actor-critic algorithms, showing that regularized policy gradient techniques can be interpreted as advantage function learning algorithms. We conclude with some numerical examples that demonstrate improved data efficiency and stability of PGQL. In particular, we tested PGQL on the full suite of Atari games and achieved performance exceeding that of both asynchronous advantage actor-critic (A3C) and Q-learning.
References in corpus (3)
Cited by in corpus (41)
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Soft Actor-Critic Algorithms and Applications
- An Introduction to Deep Reinforcement Learning
- Distributional Soft Actor-Critic: Off-Policy Reinforcement Learning for Addressing Value Estimation Errors
- Equivalence Between Policy Gradients and Soft Q-Learning
- A Survey of Deep Reinforcement Learning in Video Games
- Evolved Policy Gradients
- A unified view of entropy-regularized Markov decision processes
- Learning Montezuma's Revenge from a Single Demonstration
- Variational Inverse Control with Events: A General Framework for Data-Driven Reward Definition
- The Mirage of Action-Dependent Baselines in Reinforcement Learning
- Joint Multi-Dimension Pruning via Numerical Gradient Update
- Recall Traces: Backtracking Models for Efficient Reinforcement Learning
- Analysing Results from AI Benchmarks: Key Indicators and How to Obtain Them
- Reinforcement Learning Based Text Style Transfer without Parallel Training Corpus
- Diversity Actor-Critic: Sample-Aware Entropy Regularization for Sample-Efficient Exploration
- Effective Exploration for Deep Reinforcement Learning via Bootstrapped Q-Ensembles under Tsallis Entropy Regularization
- Path Consistency Learning in Tsallis Entropy Regularized MDPs
- Direct and indirect reinforcement learning
- Epistemic Risk-Sensitive Reinforcement Learning
- Implicit Policy for Reinforcement Learning
- Generalized Off-Policy Actor-Critic
- Off-Policy Actor-Critic in an Ensemble: Achieving Maximum General Entropy and Effective Environment Exploration in Deep Reinforcement Learning
- A Max-Min Entropy Framework for Reinforcement Learning
- Can Temporal-Difference and Q-Learning Learn Representation? A Mean-Field Theory
- Provably Convergent Two-Timescale Off-Policy Actor-Critic with Function Approximation
- P3O: Policy-on Policy-off Policy Optimization
- Soft Policy Gradient Method for Maximum Entropy Deep Reinforcement Learning
- Reinforcement Learning with Dynamic Boltzmann Softmax Updates
- A short variational proof of equivalence between policy gradients and soft Q learning
- Review, Analysis and Design of a Comprehensive Deep Reinforcement Learning Framework
- Probabilistic Mixture-of-Experts for Efficient Deep Reinforcement Learning
- Merging Deterministic Policy Gradient Estimations with Varied Bias-Variance Tradeoff for Effective Deep Reinforcement Learning
- Energy-based Surprise Minimization for Multi-Agent Value Factorization
- A Convergence Result for Regularized Actor-Critic Methods
- Using Deep Reinforcement Learning for the Continuous Control of Robotic Arms
- Measuring Progress in Deep Reinforcement Learning Sample Efficiency
- Modeling and Optimization of Epidemiological Control Policies Through Reinforcement Learning
- Weighted Entropy Modification for Soft Actor-Critic
- Ranking Policy Gradient
- Gaussian Processes for Individualized Continuous Treatment Rule Estimation