A Lyapunov Analysis of Softmax Policy Gradient for Stochastic Bandits
arXiv:2603.26547
Abstract
We adapt the analysis of policy gradient for continuous time -armed stochastic bandits by Lattimore (2026) to the standard discrete time setup. As in continuous time, we prove that with learning rate the regret is where is the horizon and and are the minimum and maximum gaps.
6 pages