A Diffusion Analysis of Policy Gradient for Stochastic Bandits
arXiv:2603.10219
Abstract
We study a continuous-time diffusion approximation of policy gradient for -armed stochastic bandits. We prove that with a learning rate the regret is where is the horizon and the minimum gap. Moreover, we construct an instance with only logarithmically many arms for which the regret is linear unless .
17 pages