On the Global Convergence of Actor-Critic: A Case for Linear Quadratic Regulator with Ergodic Cost
arXiv:1907.06246
Abstract
Despite the empirical success of the actor-critic algorithm, its theoretical understanding lags behind. In a broader context, actor-critic can be viewed as an online alternating update algorithm for bilevel optimization, whose convergence is known to be fragile. To understand the instability of actor-critic, we focus on its application to linear quadratic regulators, a simple yet fundamental setting of reinforcement learning. We establish a nonasymptotic convergence analysis of actor-critic in this setting. In particular, we prove that actor-critic finds a globally optimal pair of actor (policy) and critic (action-value function) at a linear rate of convergence. Our analysis may serve as a preliminary step towards a complete theoretical understanding of bilevel optimization with nonconvex subproblems, which is NP-hard in the worst case and is often solved using heuristics.
41 pages
References in corpus (4)
Cited by in corpus (5)
- Actor-Critic Provably Finds Nash Equilibria of Linear-Quadratic Mean-Field Games
- Combining Model-Based and Model-Free Methods for Nonlinear Control: A Provably Convergent Policy Gradient Approach
- Policy Optimization for Markovian Jump Linear Quadratic Control: Gradient-Based Methods and Global Convergence
- Analyzing the Variance of Policy Gradient Estimators for the Linear-Quadratic Regulator
- Policy Learning of MDPs with Mixed Continuous/Discrete Variables: A Case Study on Model-Free Control of Markovian Jump Systems