4 papers
Faster Fixed-Point Methods for Multichain MDPs
Matthew Zurek, Yudong Chen
We study value-iteration (VI) algorithms for solving general (a.k.a. multichain) Markov decision processes (MDPs) under the average-reward criterion, a fundamental but theoreticall…
Optimal Single-Policy Sample Complexity and Transient Coverage for Average-Reward Offline RL
Matthew Zurek, Guy Zamir, Yudong Chen
We study offline reinforcement learning in average-reward MDPs, which presents increased challenges from the perspectives of distribution shift and non-uniform coverage, and has be…
Span-Agnostic Optimal Sample Complexity and Oracle Inequalities for Average-Reward RL
Matthew Zurek, Yudong Chen
We study the sample complexity of finding an -optimal policy in average-reward Markov Decision Processes (MDPs) with a generative model. The minimax optimal span-based…
The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis
Matthew Zurek, Yudong Chen
We study the sample complexity of the plug-in approach for learning -optimal policies in average-reward Markov decision processes (MDPs) with a generative model. The p…