2 papers
math.OC2025
Policy Mirror Descent with Temporal Difference Learning: Sample Complexity under Online Markov Data
Wenye Li, Hongxu Chen, Jiacai Liu +1
This paper studies the policy mirror descent (PMD) method, which is a general policy optimization framework in reinforcement learning and can cover a wide range of policy gradient…
math.OC2025
On the Convergence of Policy Mirror Descent with Temporal Difference Evaluation
Jiacai Liu, Wenye Li, Ke Wei
Policy mirror descent (PMD) is a general policy optimization framework in reinforcement learning, which can cover a wide range of typical policy optimization methods by specifying…