26 citations · 55 across the 6 of their papers we have counts for
7 papers · 1 filter
Complexity of Finding Stationary Points of Nonsmooth Nonconvex Functions
Jingzhao Zhang, Hongzhou Lin, Stefanie Jegelka +2
We provide the first non-asymptotic analysis for finding stationary points of nonsmooth, nonconvex functions. In particular, we study the class of Hadamard semi-differentiable func…
Why are Adaptive Methods Good for Attention Models?
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit +4
While stochastic gradient descent (SGD) is still the \emph{de facto} algorithm in deep learning, adaptive methods like Clipped SGD/Adam have been observed to outperform SGD across…
Acceleration in First Order Quasi-strongly Convex Optimization by ODE Discretization
Jingzhao Zhang, Suvrit Sra, Ali Jadbabaie
We study gradient-based optimization methods obtained by direct Runge-Kutta discretization of the ordinary differential equation (ODE) describing the movement of a heavy-ball under…
Why gradient clipping accelerates training: A theoretical justification for adaptivity
Jingzhao Zhang, Tianxing He, Suvrit Sra +1
We provide a theoretical explanation for the effectiveness of gradient clipping in training deep neural networks. The key ingredient is a new smoothness condition derived from prac…
R-SPIDER: A Fast Riemannian Stochastic Optimization Algorithm with Curvature Independent Rate
Jingzhao Zhang, Hongyi Zhang, Suvrit Sra
We study smooth stochastic optimization problems on Riemannian manifolds. Via adapting the recently proposed SPIDER algorithm \citep{fang2018spider} (a variance reduced stochastic…
Achieving Acceleration in Distributed Optimization via Direct Discretization of the Heavy-Ball ODE
Jingzhao Zhang, César A. Uribe, Aryan Mokhtari +1
We develop a distributed algorithm for convex Empirical Risk Minimization, the problem of minimizing large but finite sum of convex functions over networks. The proposed algorithm…