6 papers
Internal State-Based Policy Gradient Methods for Partially Observable Markov Potential Games
Wonseok Yang, Thinh T. Doan
This letter studies multi-agent reinforcement learning in partially observable Markov potential games. Solving this problem is challenging due to partial observability, decentraliz…
Finite-Time Analysis of Projected Two-Time-Scale Stochastic Approximation
Yitao Bai, Thinh T. Doan, Justin Romberg
We study the finite-time convergence of projected linear two-time-scale stochastic approximation with constant step sizes and Polyak--Ruppert averaging. We establish an explicit me…
Fast Two-Time-Scale Stochastic Gradient Method with Applications in Reinforcement Learning
Sihan Zeng, Thinh T. Doan
Two-time-scale optimization is a framework introduced in Zeng et al. (2024) that abstracts a range of policy evaluation and policy optimization problems in reinforcement learning (…
Nonasymptotic CLT and Error Bounds for Two-Time-Scale Stochastic Approximation
Seo Taek Kong, Sihan Zeng, Thinh T. Doan +1
We consider linear two-time-scale stochastic approximation algorithms driven by martingale noise. Recent applications in machine learning motivate the need to understand finite-tim…
Accelerating Multi-Task Temporal Difference Learning under Low-Rank Representation
Yitao Bai, Sihan Zeng, Justin Romberg +1
We study policy evaluation problems in multi-task reinforcement learning (RL) under a low-rank representation setting. In this setting, we are given learning tasks where the co…
Finite-Time Complexity of Online Primal-Dual Natural Actor-Critic Algorithm for Constrained Markov Decision Processes
Sihan Zeng, Thinh T. Doan, Justin Romberg
We consider a discounted cost constrained Markov decision process (CMDP) policy optimization problem, in which an agent seeks to maximize a discounted cumulative reward subject to…