paper

On Value Iteration Convergence in Connected MDPs

arXiv:2406.09592

Abstract

This paper establishes that an MDP with a unique optimal policy and ergodic associated transition matrix ensures the convergence of various versions of the Value Iteration algorithm at a geometric rate that exceeds the discount factor γ for both discounted and average-reward criteria.

8 pages, 1 figure

On Value Iteration Convergence in Connected MDPs · wovepaper