Single-Timescale Actor-Critic Provably Finds Globally Optimal Policy
arXiv:2008.00483
Abstract
We study the global convergence and global optimality of actor-critic, one of the most popular families of reinforcement learning algorithms. While most existing works on actor-critic employ bi-level or two-timescale updates, we focus on the more practical single-timescale setting, where the actor and critic are updated simultaneously. Specifically, in each iteration, the critic update is obtained by applying the Bellman evaluation operator only once while the actor is updated in the policy gradient direction computed using the critic. Moreover, we consider two function approximation settings where both the actor and critic are represented by linear or deep neural networks. For both cases, we prove that the actor sequence converges to a globally optimal policy at a sublinear rate, where is the number of iterations. To the best of our knowledge, we establish the rate of convergence and global optimality of single-timescale actor-critic with linear function approximation for the first time. Moreover, under the broader scope of policy optimization with nonlinear function approximation, we prove that actor-critic with deep neural network finds the globally optimal policy at a sublinear rate for the first time.
References in corpus (12)
- Solving Rubik's Cube with a Robot Hand
- Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- A Theory of Regularized Markov Decision Processes
- A Two-Timescale Framework for Bilevel Optimization: Complexity Analysis and Application to Actor-Critic
- Information-Theoretic Considerations in Batch Reinforcement Learning
- Two Time-scale Off-Policy TD Learning: Non-asymptotic Analysis over Markovian Samples
- A Finite Time Analysis of Two Time-Scale Actor Critic Methods
- Non-asymptotic Convergence Analysis of Two Time-scale (Natural) Actor-Critic Algorithms
- On the Global Convergence of Actor-Critic: A Case for Linear Quadratic Regulator with Ergodic Cost
- Reinforcement Learning via Fenchel-Rockafellar Duality
- Actor-Critic Provably Finds Nash Equilibria of Linear-Quadratic Mean-Field Games
Cited by in corpus (7)
- Is Pessimism Provably Efficient for Offline RL?
- Provable Benefits of Actor-Critic Methods for Offline Reinforcement Learning
- Doubly Robust Off-Policy Actor-Critic: Convergence and Optimality
- Tighter Analysis of Alternating Stochastic Gradient Method for Stochastic Nested Problems
- Risk-Sensitive Deep RL: Variance-Constrained Actor-Critic Provably Finds Globally Optimal Policy
- Mean-Field Multi-Agent Reinforcement Learning: A Decentralized Network Approach
- On the Convergence Rate of Off-Policy Policy Optimization Methods with Density-Ratio Correction