Online Linear Quadratic Control
arXiv:1806.07104
Abstract
We study the problem of controlling linear time-invariant systems with known noisy dynamics and adversarially chosen quadratic losses. We present the first efficient online learning algorithms in this setting that guarantee regret under mild assumptions, where is the time horizon. Our algorithms rely on a novel SDP relaxation for the steady-state distribution of the system. Crucially, and in contrast to previously proposed relaxations, the feasible solutions of our SDP all correspond to "strongly stable" policies that mix exponentially fast to a steady state.
References in corpus (2)
Cited by in corpus (12)
- Logarithmic Regret Bound in Partially Observable Linear Dynamical Systems
- Online Optimization with Memory and Competitive Control
- Online Data Poisoning Attack
- The Nonstochastic Control Problem
- Dynamic Regret Minimization for Control of Non-stationary Linear Dynamical Systems
- Non-Stochastic Control with Bandit Feedback
- Learning Stabilizing Controllers for Unstable Linear Quadratic Regulators from a Single Trajectory
- Optimistic robust linear quadratic dual control
- Online Learning Robust Control of Nonlinear Dynamical Systems
- Extracting Latent State Representations with Linear Dynamics from Rich Observations
- Topological Linear System Identification via Moderate Deviations Theory
- Differentiable Robust LQR Layers