1 paper · 1 filter
Matthew Zurek, Yudong Chen
We study value-iteration (VI) algorithms for solving general (a.k.a. multichain) Markov decision processes (MDPs) under the average-reward criterion, a fundamental but theoreticall…