5 papers
Exponential Convergence of (Stochastic) Gradient Descent for Separable Logistic Regression
Sacchit Kale, Piyushi Manupriya, Pierre Marion +2
Gradient descent and stochastic gradient descent are central to modern machine learning, yet their behavior under large step sizes remains theoretically unclear. Recent work sugges…
On the Generalization and Robustness in Conditional Value-at-Risk
Dinesh Karthik Mulumudi, Piyushi Manupriya, Gholamali Aminian +1
Conditional Value-at-Risk (CVaR) is a widely used risk-sensitive objective for learning under rare but high-impact losses, yet its statistical behavior under heavy-tailed data rema…
Beyond Propagation of Chaos: A Stochastic Algorithm for Mean Field Optimization
Chandan Tankala, Dheeraj M. Nagaraj, Anant Raj
Gradient flow in the 2-Wasserstein space is widely used to optimize functionals over probability distributions and is typically implemented using an interacting particle system wit…
Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates
Jincheng Mei, Bo Dai, Alekh Agarwal +4
We provide a new understanding of the stochastic gradient bandit algorithm by showing that it converges to a globally optimal policy almost surely using \emph{any} constant learnin…
Towards Principled, Practical Policy Gradient for Bandits and Tabular MDPs
Michael Lu, Matin Aghaei, Anant Raj +1
We consider (stochastic) softmax policy gradient (PG) methods for bandits and tabular Markov decision processes (MDPs). While the PG objective is non-concave, recent research has u…