8 citations · 42 across the 20 of their papers we have counts for
7 papers · 1 filter
Implicit Regularization and Convergence for Weight Normalization
Xiaoxia Wu, Edgar Dobriban, Tongzheng Ren +5
Normalization methods such as batch [Ioffe and Szegedy, 2015], weight [Salimansand Kingma, 2016], instance [Ulyanov et al., 2016], and layer normalization [Baet al., 2016] have bee…
Faster Johnson-Lindenstrauss Transforms via Kronecker Products
Ruhui Jin, Tamara G. Kolda, Rachel Ward
The Kronecker product is an important matrix operation with a wide range of applications in supporting fast linear transforms, including signal processing, graph theory, quantum co…
Linear Convergence of Adaptive Stochastic Gradient Descent
Yuege Xie, Xiaoxia Wu, Rachel Ward
We prove that the norm version of the adaptive stochastic gradient method (AdaGrad-Norm) achieves a linear convergence rate for a subset of either strongly convex functions or non-…
Bias of Homotopic Gradient Descent for the Hinge Loss
Denali Molitor, Deanna Needell, Rachel Ward
Gradient descent is a simple and widely used optimization method for machine learning. For homogeneous linear classifiers applied to separable data, gradient descent has been shown…
Concentration inequalities for random matrix products
Amelia Henriksen, Rachel Ward
Suppose is a sequence of bounded independent random matrices with common dimension and common expectation . Under t…
AdaOja: Adaptive Learning Rates for Streaming PCA
Amelia Henriksen, Rachel Ward
Oja's algorithm has been the cornerstone of streaming methods in Principal Component Analysis (PCA) since it was first proposed in 1982. However, Oja's algorithm does not have a st…