Equivalence Between Policy Gradients and Soft Q-Learning
arXiv:1704.06440
Abstract
Two of the leading approaches for model-free reinforcement learning are policy gradient methods and -learning methods. -learning methods can be effective and sample-efficient when they work, however, it is not well-understood why they work, since empirically, the -values they estimate are very inaccurate. A partial explanation may be that -learning methods are secretly implementing policy gradient updates: we show that there is a precise equivalence between -learning and policy gradient methods in the setting of entropy-regularized reinforcement learning, that "soft" (entropy-regularized) -learning is exactly equivalent to a policy gradient method. We also point out a connection between -learning methods and natural policy gradient methods. Experimentally, we explore the entropy-regularized versions of -learning and policy gradients, and we find them to perform as well as (or slightly better than) the standard variants on the Atari benchmark. We also show that the equivalence holds in practical settings by constructing a -learning method that closely matches the learning dynamics of A3C without using a target network or -greedy exploration schedule.
Cited by in corpus (97)
- A Brief Survey of Deep Reinforcement Learning
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Soft Actor-Critic Algorithms and Applications
- An Introduction to Deep Reinforcement Learning
- A Survey and Critique of Multiagent Deep Reinforcement Learning
- Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
- Distributional Soft Actor-Critic: Off-Policy Reinforcement Learning for Addressing Value Estimation Errors
- Maximum a Posteriori Policy Optimisation
- Multi-Goal Reinforcement Learning: Challenging Robotics Environments and Request for Research
- Distral: Robust Multitask Reinforcement Learning
- Transfer Learning in Deep Reinforcement Learning: A Survey
- A Survey of Deep Reinforcement Learning in Video Games
- Evolved Policy Gradients
- Reinforcement Learning from Imperfect Demonstrations
- Learning Montezuma's Revenge from a Single Demonstration
- Self-Consistent Trajectory Autoencoder: Hierarchical Reinforcement Learning with Trajectory Embeddings
- Latent Space Policies for Hierarchical Reinforcement Learning
- Discrete Sequential Prediction of Continuous Actions for Deep RL
- Towards Characterizing Divergence in Deep Q-Learning
- Variational Inverse Control with Events: A General Framework for Data-Driven Reward Definition
- First Order Constrained Optimization in Policy Space
- Learning Implicit Credit Assignment for Cooperative Multi-Agent Reinforcement Learning
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement Learning
- Understanding the impact of entropy on policy optimization
- Robust Reinforcement Learning for Continuous Control with Model Misspecification
- Fast Global Convergence of Natural Policy Gradient Methods with Entropy Regularization
- Distilling Policy Distillation
- VIREL: A Variational Inference Framework for Reinforcement Learning
- Trust-PCL: An Off-Policy Trust Region Method for Continuous Control
- Action and Perception as Divergence Minimization
- Hierarchical Reinforcement Learning via Advantage-Weighted Information Maximization
- Tsallis Reinforcement Learning: A Unified Framework for Maximum Entropy Reinforcement Learning
- Exploiting Hierarchy for Learning and Transfer in KL-regularized RL
- Average-reward model-free reinforcement learning: a systematic review and literature mapping
- An Information-Theoretic Optimality Principle for Deep Reinforcement Learning
- A Unified Bellman Optimality Principle Combining Reward Maximization and Empowerment
- Efficient (Soft) Q-Learning for Text Generation with Limited Good Data
- Behavior Priors for Efficient Reinforcement Learning
- Variational Bayesian Reinforcement Learning with Regret Bounds
- Neural Temporal-Difference and Q-Learning Provably Converge to Global Optima
- Connecting the Dots Between MLE and RL for Sequence Prediction
- A Tutorial on Sparse Gaussian Processes and Variational Inference
- Soft Hindsight Experience Replay
- DiCE: The Infinitely Differentiable Monte-Carlo Estimator
- Diversity Actor-Critic: Sample-Aware Entropy Regularization for Sample-Efficient Exploration
- Greedification Operators for Policy Optimization: Investigating Forward and Reverse KL Divergences
- Direct and indirect reinforcement learning
- Doubly Robust Off-Policy Actor-Critic: Convergence and Optimality
- Effective Exploration for Deep Reinforcement Learning via Bootstrapped Q-Ensembles under Tsallis Entropy Regularization
- Marginalized State Distribution Entropy Regularization in Policy Optimization
- Entropy Regularization with Discounted Future State Distribution in Policy Gradient Methods
- SOAC: The Soft Option Actor-Critic Architecture
- Local Search for Policy Iteration in Continuous Control
- Improving Computational Efficiency in Visual Reinforcement Learning via Stored Embeddings
- A Regularized Opponent Model with Maximum Entropy Objective
- Information Theoretic Model Predictive Q-Learning
- Implicit Policy for Reinforcement Learning
- Generalized Off-Policy Actor-Critic
- Provably Convergent Two-Timescale Off-Policy Actor-Critic with Function Approximation
- Can Temporal-Difference and Q-Learning Learn Representation? A Mean-Field Theory
- An operator view of policy gradient methods
- A Max-Min Entropy Framework for Reinforcement Learning
- Steady State Analysis of Episodic Reinforcement Learning
- Planning in entropy-regularized Markov decision processes and games
- Soft Policy Gradient Method for Maximum Entropy Deep Reinforcement Learning
- RL STaR Platform: Reinforcement Learning for Simulation based Training of Robots
- Inverse Reinforcement Learning from a Gradient-based Learner
- Review, Analysis and Design of a Comprehensive Deep Reinforcement Learning Framework
- A short variational proof of equivalence between policy gradients and soft Q learning
- Mutual-Information Regularization in Markov Decision Processes and Actor-Critic Learning
- GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning
- Revisiting Prioritized Experience Replay: A Value Perspective
- From Credit Assignment to Entropy Regularization: Two New Algorithms for Neural Sequence Prediction
- Pretrain Soft Q-Learning with Imperfect Demonstrations
- Exploration by Maximizing Rényi Entropy for Reward-Free RL Framework
- Approximate Inference in Discrete Distributions with Monte Carlo Tree Search and Value Functions
- Variational Transport: A Convergent Particle-BasedAlgorithm for Distributional Optimization
- Finding the Near Optimal Policy via Adaptive Reduced Regularization in MDPs
- Estimating Optimal Infinite Horizon Dynamic Treatment Regimes via pT-Learning
- OPAC: Opportunistic Actor-Critic
- A Convergence Result for Regularized Actor-Critic Methods
- Energy-based Surprise Minimization for Multi-Agent Value Factorization
- Improved Soft Actor-Critic: Mixing Prioritized Off-Policy Samples with On-Policy Experience
- A Relation Analysis of Markov Decision Process Frameworks
- Reinforcement Learning as Iterative and Amortised Inference
- Learning Policies through Quantile Regression
- An Entropy Regularization Free Mechanism for Policy-based Reinforcement Learning
- Will it Blend? Composing Value Functions in Reinforcement Learning
- Advanced Policies: A First-Principles Path from Policy Gradient to Q-Learning
- Regularized Policies are Reward Robust
- Convex Regularization in Monte-Carlo Tree Search
- Distributionally-Constrained Policy Optimization via Unbalanced Optimal Transport
- Weighted Entropy Modification for Soft Actor-Critic
- Gaussian Processes for Individualized Continuous Treatment Rule Estimation
- Regularize! Don't Mix: Multi-Agent Reinforcement Learning without Explicit Centralized Structures
- Generative Actor-Critic: An Off-policy Algorithm Using the Push-forward Model
- Towards an Understanding of Default Policies in Multitask Policy Optimization