2 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.LG2024
Stochastic Gradient Succeeds for Bandits
Jincheng Mei, Zixin Zhong, Bo Dai +3
We show that the \emph{stochastic gradient} bandit algorithm converges to a \emph{globally optimal} policy at an rate, even with a \emph{constant} step size. Remarkably, g…
cs.LG2023
Regularization and Variance-Weighted Regression Achieves Minimax Optimality in Linear MDPs: Theory and Practice
Toshinori Kitamura, Tadashi Kozuno, Yunhao Tang +12
Mirror descent value iteration (MDVI), an abstraction of Kullback-Leibler (KL) and entropy-regularized reinforcement learning (RL), has served as the basis for recent high-performi…
cs.LG2023★ 2 cited
The Role of Baselines in Policy Gradient Optimization
Jincheng Mei, Wesley Chung, Valentin Thomas +3
We study the effect of baselines in on-policy stochastic policy gradient optimization, and close the gap between the theory and practice of policy optimization methods. Our first c…