Bootstrap Policy Iteration for Stochastic Linear Quadratic Tracking with Multiplicative Noise
arXiv:2508.20394
Abstract
This paper studies the linear quadratic tracking problem for continuous-time stochastic systems with multiplicative noise. The proposed framework formulates the problem under an average cost criterion and separates the computation of the optimal feedback and feedforward gains. By developing a bootstrap policy iteration algorithm, we eliminate the restrictive a priori requirement for an initial mean square stabilizing feedback gain in existing policy iteration methods. Based on this iterative framework, an off-policy reinforcement learning algorithm is proposed to learn the optimal feedback gain directly from data. Using the learned feedback gain, the feedforward gain is subsequently obtained through a data-driven one-shot computation procedure. These components work together to provide a model-free solution to the stochastic optimal tracking control problem. The effectiveness of the proposed method is demonstrated through a numerical example.