42 citations · 140 across the 33 of their papers we have counts for
Showing stat.MLShow all
3 papers · 1 filter
stat.ML2024
Global Convergence in Training Large-Scale Transformers
Cheng Gao, Yuan Cao, Zihao Li +5
Despite the widespread success of Transformers across various domains, their optimization guarantees in large-scale model settings are not well-understood. This paper rigorously an…
stat.ML2016★ 42 cited
Stochastic Primal-Dual Methods and Sample Complexity of Reinforcement Learning
Yichen Chen, Mengdi Wang
We study the online estimation of the optimal policy of a Markov decision process (MDP). We propose a class of Stochastic Primal-Dual (SPD) methods which exploit the inherent minim…
stat.ML2014★ 6 cited
Stochastic Compositional Gradient Descent: Algorithms for Minimizing Compositions of Expected-Value Functions
Mengdi Wang, Ethan X. Fang, Han Liu
Classical stochastic gradient methods are well suited for minimizing expected-value objective functions. However, they do not apply to the minimization of a nonlinear function invo…