Naive Exploration is Optimal for Online LQR
arXiv:2001.09576
Abstract
We consider the problem of online adaptive control of the linear quadratic regulator, where the true system parameters are unknown. We prove new upper and lower bounds demonstrating that the optimal regret scales as , where is the number of time steps, is the dimension of the input space, and is the dimension of the system state. Notably, our lower bounds rule out the possibility of a -regret algorithm, which had been conjectured due to the apparent strong convexity of the problem. Our upper bound is attained by a simple variant of , where the learner selects control inputs according to the optimal controller for their estimate of the system while injecting exploratory random noise. While this approach was shown to achieve -regret by (Mania et al. 2019), we show that if the learner continually refines their estimates of the system matrices, the method attains optimal dimension dependence as well. Central to our upper and lower bounds is a new approach for controlling perturbations of Riccati equations called the , which we use to derive suboptimality bounds for the certainty equivalent controller synthesized from estimated system dynamics. This in turn enables regret upper bounds which hold for and scale with natural control-theoretic quantities.
References in corpus (4)
Cited by in corpus (23)
- Improper Learning for Non-Stochastic Control
- Logarithmic Regret Bound in Partially Observable Linear Dynamical Systems
- Information Theoretic Regret Bounds for Online Nonlinear Control
- Learning nonlinear dynamical systems from a single trajectory
- Black-Box Control for Linear Dynamical Systems
- Logarithmic Regret for Learning Linear Quadratic Regulators Efficiently
- Combining Model-Based and Model-Free Methods for Nonlinear Control: A Provably Convergent Policy Gradient Approach
- The Power of Predictions in Online Control
- Rebounding Bandits for Modeling Satiation Effects
- Making Non-Stochastic Control (Almost) as Easy as Stochastic
- Patterns, predictions, and actions: A story about machine learning
- Dynamic Regret Minimization for Control of Non-stationary Linear Dynamical Systems
- Certainty Equivalent Perception-Based Control
- Learning Stabilizing Controllers for Unstable Linear Quadratic Regulators from a Single Trajectory
- Bandit Linear Control
- SLIP: Learning to Predict in Unknown Dynamical Systems with Long-Term Memory
- Thompson sampling for linear quadratic mean-field teams
- Task-Optimal Exploration in Linear Dynamical Systems
- On Uninformative Optimal Policies in Adaptive LQR with Unknown B-Matrix
- How Are Learned Perception-Based Controllers Impacted by the Limits of Robust Control?
- Certainty Equivalent Quadratic Control for Markov Jump Systems
- Geometric Exploration for Online Control
- Exact Asymptotics for Linear Quadratic Adaptive Control