A Theory of Regularized Markov Decision Processes
arXiv:1901.11275
Abstract
Many recent successful (deep) reinforcement learning algorithms make use of regularization, generally based on entropy or Kullback-Leibler divergence. We propose a general theory of regularized Markov Decision Processes that generalizes these approaches in two directions: we consider a larger class of regularizers, and we consider the general modified policy iteration approach, encompassing both policy iteration and value iteration. The core building blocks of this theory are a notion of regularized Bellman operator and the Legendre-Fenchel transform, a classical tool of convex optimization. This approach allows for error propagation analyses of general algorithmic schemes of which (possibly variants of) classical algorithms such as Trust Region Policy Optimization, Soft Q-learning, Stochastic Actor Critic or Dynamic Policy Programming are special cases. This also draws connections to proximal convex optimization, especially to Mirror Descent.
ICML 2019
Cited by in corpus (13)
- From Poincaré Recurrence to Convergence in Imperfect Information Games: Finding Equilibrium via Regularization
- Provable Benefits of Actor-Critic Methods for Offline Reinforcement Learning
- Adversarially Guided Actor-Critic
- MADE: Exploration via Maximizing Deviation from Explored Regions
- Offline Reinforcement Learning with Value-based Episodic Memory
- On Connections between Constrained Optimization and Reinforcement Learning
- Twice regularized MDPs and the equivalence between robustness and regularization
- Dimension-Free Rates for Natural Policy Gradient in Multi-Agent Reinforcement Learning
- Provably Correct Optimization and Exploration with Non-linear Policies
- On the Convergence of Approximate and Regularized Policy Iteration Schemes
- Robust Generalization despite Distribution Shift via Minimum Discriminating Information
- Geometric Value Iteration: Dynamic Error-Aware KL Regularization for Reinforcement Learning
- Cautious Actor-Critic