optimal control

Model-Free Q-Learning for Infinite-Horizon Stochastic Linear Quadratic Problems with Regime Switching

arXiv:2607.26456

summary

The paper develops model‑free Q‑learning algorithms that learn optimal controllers for infinite‑horizon continuous‑time stochastic linear‑quadratic problems with regime switching, using only online state data.

Abstract

This paper addresses infinite-horizon continuous-time stochastic linear quadratic optimal control problems with regime switching. We propose a paradigm shift from model-based design by adopting an adaptive dynamic programming approach, specifically developing on-policy and off-policy Q-learning algorithms that learn the optimal controller solely from online state trajectory data. The theoretical core of our work consists of a complete proof of the equivalence between the on- and off-policy architectures, alongside a rigorous analysis establishing the stability of the closed-loop system and the convergence of the algorithms to the optimal solution. For computational tractability, we implement these algorithms using vectorization and Kronecker product algebra. The theoretical results are corroborated by numerical case studies that clearly demonstrate the operational effectiveness and practical feasibility of the proposed model-free control strategy.

21 pages, 1 figure

Topics & keywords

#stochastic control#linear quadratic#regime switching#reinforcement learning#q-learninginfinite-horizoncontinuous-timeadaptive dynamic programmingon-policyoff-policyKronecker product
Model-Free Q-Learning for Infinite-Horizon Stochastic Linear Quadratic Problems with Regime Switching · wovepaper