paper

The Sample Complexity of Policy Learning with Mu-Resets

arXiv:2608.07772

Abstract

We study policy-based reinforcement learning under the -resets interaction protocol of Kakade and Langford [KL02]. This interaction protocol enables the learner to sample trajectories from a given exploratory reset distribution , in addition to the starting distribution. We resolve the question raised by [KLS25] on the role of policy realizability for the sample complexity of this problem. Critically, the dependence on horizon is governed by the notion of coverage assumed of the reset distribution. Under bounded all-policy concentrability, we show a sample complexity lower bound; with bounded pushforward concentrability, we show the dependence on horizon is tightly characterized as .

comments welcome

The Sample Complexity of Policy Learning with Mu-Resets · wovepaper