3 papers
cs.LG2025
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
Max Qiushi Lin, Jincheng Mei, Matin Aghaei +6
Policy gradient (PG) methods have played an essential role in the empirical successes of reinforcement learning. In order to handle large state-action spaces, PG methods are typica…
cs.LG2025
Ordering-based Conditions for Global Convergence of Policy Gradient Methods
Jincheng Mei, Bo Dai, Alekh Agarwal +3
We prove that, for finite-arm bandits with linear function approximation, the global convergence of policy gradient (PG) methods depends on inter-related properties between the pol…
cs.LG2025
Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates
Jincheng Mei, Bo Dai, Alekh Agarwal +4
We provide a new understanding of the stochastic gradient bandit algorithm by showing that it converges to a globally optimal policy almost surely using \emph{any} constant learnin…