4 papers · 1 filter
On the Policy Convergence of Policy Mirror Descent Methods
Wenye Li, Ke Wei
We study the policy convergence of unregularized policy mirror descent (PMD) with arbitrary constant step sizes for finite discounted Markov decision processes. We focus on decompo…
Policy Mirror Descent with Temporal Difference Learning: Sample Complexity under Online Markov Data
Wenye Li, Hongxu Chen, Jiacai Liu +1
This paper studies the policy mirror descent (PMD) method, which is a general policy optimization framework in reinforcement learning and can cover a wide range of policy gradient…
On the Convergence of Policy Mirror Descent with Temporal Difference Evaluation
Jiacai Liu, Wenye Li, Ke Wei
Policy mirror descent (PMD) is a general policy optimization framework in reinforcement learning, which can cover a wide range of typical policy optimization methods by specifying…
Elementary Analysis of Policy Gradient Methods
Jiacai Liu, Wenye Li, Ke Wei
Projected policy gradient under the simplex parameterization, policy gradient and natural policy gradient under the softmax parameterization, are fundamental algorithms in reinforc…