paper

Exponential Lower Bounds For Policy Iteration

arXiv:1003.3418

Abstract

We study policy iteration for infinite-horizon Markov decision processes. It has recently been shown policy iteration style algorithms have exponential lower bounds in a two player game setting. We extend these lower bounds to Markov decision processes with the total reward and average-reward optimality criteria.

References in corpus (1)

Exponential Lower Bounds For Policy Iteration · wovepaper