The Sample Complexity of Policy Learning with Mu-Resets
arXiv:2608.07772
Abstract
We study policy-based reinforcement learning under the -resets interaction protocol of Kakade and Langford [KL02]. This interaction protocol enables the learner to sample trajectories from a given exploratory reset distribution , in addition to the starting distribution. We resolve the question raised by [KLS25] on the role of policy realizability for the sample complexity of this problem. Critically, the dependence on horizon is governed by the notion of coverage assumed of the reset distribution. Under bounded all-policy concentrability, we show a sample complexity lower bound; with bounded pushforward concentrability, we show the dependence on horizon is tightly characterized as .
comments welcome