3 papers
math.OC2026
Value Mirror Descent for Reinforcement Learning
Zhichao Jia, Guanghui Lan
Value iteration-type methods have been extensively studied for computing a nearly optimal value function in reinforcement learning (RL). Under a generative sampling model, these me…
cs.LG2026
Actor-Accelerated Policy Dual Averaging for Reinforcement Learning in Continuous Action Spaces
Ji Gao, Caleb Ju, Guanghui Lan +1
Policy Dual Averaging (PDA) offers a principled Policy Mirror Descent (PMD) framework that more naturally admits value function approximation than standard PMD, enabling the use of…
cs.LG2024
Strongly-polynomial time and validation analysis of policy gradient methods
Caleb Ju, Guanghui Lan
This paper proposes a novel termination criterion, termed the advantage gap function, for finite state and action Markov decision processes (MDP) and reinforcement learning (RL). B…