Model-Free Linear Quadratic Control via Reduction to Expert Prediction
arXiv:1804.06021
Abstract
Model-free approaches for reinforcement learning (RL) and continuous control find policies based only on past states and rewards, without fitting a model of the system dynamics. They are appealing as they are general purpose and easy to implement; however, they also come with fewer theoretical guarantees than model-based RL. In this work, we present a new model-free algorithm for controlling linear quadratic (LQ) systems, and show that its regret scales as for any small if time horizon satisfies for a constant . The algorithm is based on a reduction of control of Markov decision processes to an expert prediction problem. In practice, it corresponds to a variant of policy iteration with forced exploration, where the policy in each phase is greedy with respect to the average of all previous value functions. This is the first model-free algorithm for adaptive control of LQ systems that provably achieves sublinear regret and has a polynomial computation cost. Empirically, our algorithm dramatically outperforms standard policy iteration, but performs worse than a model-based approach.
Cited by in corpus (28)
- Logarithmic Regret Bound in Partially Observable Linear Dynamical Systems
- Certainty Equivalence is Efficient for Linear Quadratic Control
- Variance-reduced -learning is minimax optimal
- Finite-time Analysis of Approximate Policy Iteration for the Linear Quadratic Regulator
- Exploration-Enhanced POLITEX
- Average-reward model-free reinforcement learning: a systematic review and literature mapping
- From self-tuning regulators to reinforcement learning and back again
- Adaptive Regret for Control of Time-Varying Dynamics
- Logarithmic Regret for Learning Linear Quadratic Regulators Efficiently
- The Nonstochastic Control Problem
- Adaptive Approximate Policy Iteration
- Efficient Learning of Distributed Linear-Quadratic Controllers
- Regret Minimization in Partially Observable Linear Quadratic Control
- Model-free optimal control of discrete-time systems with additive and multiplicative noises
- Non-Stochastic Control with Bandit Feedback
- Robust Policy Iteration for Continuous-time Linear Quadratic Regulation
- Certainty Equivalent Perception-Based Control
- Robust Reinforcement Learning: A Case Study in Linear Quadratic Regulation
- Towards a Dimension-Free Understanding of Adaptive Linear Control
- Learning Partially Observed Linear Dynamical Systems from Logarithmic Number of Samples
- On Uninformative Optimal Policies in Adaptive LQR with Unknown B-Matrix
- Online Policy Gradient for Model Free Learning of Linear Quadratic Regulators with Regret
- Continuous Control with Contexts, Provably
- Average Cost Optimal Control of Stochastic Systems Using Reinforcement Learning
- Identification and Adaptive Control of Markov Jump Systems: Sample Complexity and Regret Bounds
- Approximate Midpoint Policy Iteration for Linear Quadratic Control
- Alice's Adventures in the Markovian World
- Safe non-smooth black-box optimization with application to policy search