4 papers
Achieving Dependence for Average-Reward Q-Learning with a New Contraction Principle
Zijun Chen, Zaiwei Chen, Nian Si +1
We present the convergence rates of synchronous and asynchronous Q-learning for average-reward Markov decision processes, where the absence of contraction poses a fundamental chall…
Bellman Optimality of Average-Reward Robust Markov Decision Processes with a Constant Gain
Shengbo Wang, Nian Si
Learning and optimal control under robust Markov decision processes (MDPs) have received increasing attention, yet most existing theory, algorithms, and applications focus on finit…
Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning
Zijun Chen, Shengbo Wang, Nian Si
Motivated by practical applications where stable long-term performance is critical-such as robotics, operations research, and healthcare-we study the problem of distributionally ro…
Tractable Robust Markov Decision Processes
Julien Grand-Clément, Nian Si, Shengbo Wang
In this paper we investigate the tractability of robust Markov Decision Processes (RMDPs) under various structural assumptions on the uncertainty set. Surprisingly, we show that in…