4 papers
Average-reward reinforcement learning in semi-Markov decision processes via relative value iteration
Huizhen Yu, Yi Wan, Richard S. Sutton
This paper applies the authors' recent results on asynchronous stochastic approximation (SA) in the Borkar-Meyn framework to reinforcement learning in average-reward semi-Markov de…
Asynchronous Stochastic Approximation with Applications to Average-Reward Reinforcement Learning
Huizhen Yu, Yi Wan, Richard S. Sutton
This paper investigates the stability and convergence properties of asynchronous stochastic approximation (SA) algorithms, with a focus on extensions relevant to average-reward rei…
On Convergence of Average-Reward Q-Learning in Weakly Communicating Markov Decision Processes
Yi Wan, Huizhen Yu, Richard S. Sutton
This paper analyzes reinforcement learning (RL) algorithms for Markov decision processes (MDPs) under the average-reward criterion. We focus on Q-learning algorithms based on relat…
Reward Centering
Abhishek Naik, Yi Wan, Manan Tomar +1
We show that discounted methods for solving continuing reinforcement learning problems can perform significantly better if they center their rewards by subtracting out the rewards'…