1 citations · 1 across the 9 of their papers we have counts for
14 papers · 1 filter
Convergence of Steepest Descent and Adam under Non-Uniform Smoothness
Sharan Vaswani, Yifan Sun, Reza Babanezhad
Recent work has analyzed the convergence of first-order methods under non-uniform smoothness assumptions that better model the loss landscape in machine learning tasks. We generali…
Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs
Michael Lu, Max Qiushi Lin, Mo Chen +1
We study policy optimization for infinite-horizon, discounted constrained Markov decision processes (CMDPs). While existing theoretical guarantees typically hold for the mixture po…
Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives
Reza Asad, Reza Babanezhad, Sharan Vaswani
While Soft Actor-Critic (SAC) is highly effective in continuous control, its discrete counterpart (DSAC) performs poorly on challenging discrete-action domains such as Atari. Conse…
Optimistic Actor-Critic with Parametric Policies for Linear Markov Decision Processes
Max Qiushi Lin, Reza Asad, Kevin Tan +3
Although actor-critic methods have been successful in practice, their theoretical analyses have several limitations. Specifically, existing theoretical work either sidesteps the ex…
Towards Parameter-Free Temporal Difference Learning
Yunxiang Li, Mark Schmidt, Reza Babanezhad +1
Temporal difference (TD) learning is a fundamental algorithm for estimating value functions in reinforcement learning. Recent finite-time analyses of TD with linear function approx…
Armijo Line-search Can Make (Stochastic) Gradient Descent Provably Faster
Sharan Vaswani, Reza Babanezhad
Armijo line-search (Armijo-LS) is a standard method to set the step-size for gradient descent (GD). For smooth functions, Armijo-LS alleviates the need to know the global smoothnes…