4 citations · 4 across the 2 of their papers we have counts for
2 papers
cs.LG2024
On Convergence of Average-Reward Q-Learning in Weakly Communicating Markov Decision Processes
Yi Wan, Huizhen Yu, Richard S. Sutton
This paper analyzes reinforcement learning (RL) algorithms for Markov decision processes (MDPs) under the average-reward criterion. We focus on Q-learning algorithms based on relat…
math.OC2014★ 4 cited
Stochastic Shortest Path Games and Q-Learning
Huizhen Yu
We consider a class of two-player zero-sum stochastic games with finite state and compact control spaces, which we call stochastic shortest path (SSP) games. They are undiscounted…