paper

A Diffusion Analysis of Policy Gradient for Stochastic Bandits

arXiv:2603.10219

Abstract

We study a continuous-time diffusion approximation of policy gradient for -armed stochastic bandits. We prove that with a learning rate the regret is where is the horizon and the minimum gap. Moreover, we construct an instance with only logarithmically many arms for which the regret is linear unless .

17 pages