paper

A Lyapunov Analysis of Softmax Policy Gradient for Stochastic Bandits

arXiv:2603.26547

Abstract

We adapt the analysis of policy gradient for continuous time -armed stochastic bandits by Lattimore (2026) to the standard discrete time setup. As in continuous time, we prove that with learning rate the regret is where is the horizon and and are the minimum and maximum gaps.

6 pages