1 paper · 1 filter
Tor Lattimore
We study a continuous-time diffusion approximation of policy gradient for k-armed stochastic bandits. We prove that with a learning rate I^⋅=O(I^2/log(n)) the regret is $O(k…