2 papers
cs.LG2025
Average-reward reinforcement learning in semi-Markov decision processes via relative value iteration
Huizhen Yu, Yi Wan, Richard S. Sutton
This paper applies the authors' recent results on asynchronous stochastic approximation (SA) in the Borkar-Meyn framework to reinforcement learning in average-reward semi-Markov de…
cs.LG2024
On Convergence of Average-Reward Q-Learning in Weakly Communicating Markov Decision Processes
Yi Wan, Huizhen Yu, Richard S. Sutton
This paper analyzes reinforcement learning (RL) algorithms for Markov decision processes (MDPs) under the average-reward criterion. We focus on Q-learning algorithms based on relat…