7 citations · 7 across the 3 of their papers we have counts for
3 papers
cs.LG2024
Sample-efficient Learning of Infinite-horizon Average-reward MDPs with General Function Approximation
Jianliang He, Han Zhong, Zhuoran Yang
We study infinite-horizon average-reward Markov decision processes (AMDPs) in the context of general function approximation. Specifically, we propose a novel algorithmic framework…
cs.LG2023
Posterior Sampling for Competitive RL: Function Approximation and Partial Observation
Shuang Qiu, Ziyu Dai, Han Zhong +3
This paper investigates posterior sampling algorithms for competitive reinforcement learning (RL) in the context of general function approximations. Focusing on zero-sum Markov gam…
cs.LG2023★ 7 cited
A Theoretical Analysis of Optimistic Proximal Policy Optimization in Linear Markov Decision Processes
Han Zhong, Tong Zhang
The proximal policy optimization (PPO) algorithm stands as one of the most prosperous methods in the field of reinforcement learning (RL). Despite its success, the theoretical unde…