Strategies in POMDPs with Stage Duration
arXiv:2603.16055
Abstract
Partially observable Markov decision processes (POMDPs) with stage duration provide a framework for approximating continuous-time behavior by scaling transition probabilities with a stage duration parameter . While previous literature has primarily focused on the limit of the discounted value as the stage duration vanishes, this paper investigates the global behavior of the asymptotic value, , across varying stage durations. Our main result demonstrates that any strategy in a POMDP with stage duration can be mimicked in the base POMDP (). Specifically, we provide an explicit construction showing that for any strategy in the POMDP with stage duration , there exists a strategy in the base POMDP that secures the same asymptotic payoff. As a consequence of this theorem, we establish that the value function is nondecreasing with respect to , and that the continuous-time limit exists.