Distributionally Robust Regret Optimal LQR with Common Stage-Law Ambiguity
arXiv:2604.06158
Abstract
We study what is, to our knowledge, the first tractable multistage ex-ante distributionally robust regret optimization (DRRO) formulation for stochastic control. We consider finite-horizon LQR with common stage-law ambiguity, where disturbances are independent across time but drawn from the same unknown stage law whose mean and covariance lie in a Gelbrich ball around nominal moments. Unlike the benign single-stage quadratic setting, the nominal controller is generally not regret-optimal: reuse of the stage law makes past disturbances informative for future decisions. Despite the general hardness of DRRO, we show that, over affine disturbance-feedback policies, the multistage DRRO-LQR problem admits an exact semidefinite programming reformulation. An optimal controller in this class is the nominal LQR controller plus a strictly causal empirical-mean correction. We also characterize worst-case moment pairs and show that, for the DRRO-optimal policy, they are not unique. Portfolio liquidation experiments show that DRRO substantially reduces worst-case regret relative to DRO and the nominal controller, with comparatively modest increases in worst-case cost, and exhibits a learning effect: its correction matrices empirically approach the corresponding coefficients of the oracle controller that knows the true disturbance law in hindsight.
16 pages, 3 figures. A version of this paper has been accepted for publication in the proceedings of the 65th IEEE Conference on Decision and Control (CDC 2026)