Showing math.OCShow all
2 papers · 1 filter
math.OC2026
Value Mirror Descent for Reinforcement Learning
Zhichao Jia, Guanghui Lan
Value iteration-type methods have been extensively studied for computing a nearly optimal value function in reinforcement learning (RL). Under a generative sampling model, these me…
math.OC2025
A model-free first-order method for linear quadratic regulator with sampling complexity
Caleb Ju, Georgios Kotsalis, Guanghui Lan
We consider the classic stochastic linear quadratic regulator (LQR) problem under an infinite horizon average stage cost. By leveraging recent policy gradient methods from reinforc…