works on

From the 1 of 7 linked papers with an AI index.

activity
20242026
collaborators
Showing math.OCShow all

5 papers · 1 filter

math.OC2026

On the Policy Convergence of Policy Mirror Descent Methods

Wenye Li, Ke Wei

The paper provides a unified convergence analysis for unregularized policy mirror descent with constant step sizes in finite discounted Markov decision processes, covering a wide r…

math.OC2025

Policy Mirror Descent with Temporal Difference Learning: Sample Complexity under Online Markov Data

Wenye Li, Hongxu Chen, Jiacai Liu +1

This paper studies the policy mirror descent (PMD) method, which is a general policy optimization framework in reinforcement learning and can cover a wide range of policy gradient…

math.OC2025

On the Convergence of Policy Mirror Descent with Temporal Difference Evaluation

Jiacai Liu, Wenye Li, Ke Wei

Policy mirror descent (PMD) is a general policy optimization framework in reinforcement learning, which can cover a wide range of typical policy optimization methods by specifying…

math.OC2024

On the Convergence of Projected Policy Gradient for Any Constant Step Sizes

Jiacai Liu, Wenye Li, Dachao Lin +2

Projected policy gradient (PPG) is a basic policy optimization method in reinforcement learning. Given access to exact policy evaluations, previous studies have established the sub…

math.OC2024

Elementary Analysis of Policy Gradient Methods

Jiacai Liu, Wenye Li, Ke Wei

Projected policy gradient under the simplex parameterization, policy gradient and natural policy gradient under the softmax parameterization, are fundamental algorithms in reinforc…