Implicit Regularization and Convergence for Weight Normalization
arXiv:1911.07956
Abstract
Normalization methods such as batch [Ioffe and Szegedy, 2015], weight [Salimansand Kingma, 2016], instance [Ulyanov et al., 2016], and layer normalization [Baet al., 2016] have been widely used in modern machine learning. Here, we study the weight normalization (WN) method [Salimans and Kingma, 2016] and a variant called reparametrized projected gradient descent (rPGD) for overparametrized least-squares regression. WN and rPGD reparametrize the weights with a scale g and a unit vector w and thus the objective function becomes non-convex. We show that this non-convex formulation has beneficial regularization effects compared to gradient descent on the original objective. These methods adaptively regularize the weights and converge close to the minimum l2 norm solution, even for initializations far from zero. For certain stepsizes of g and w , we show that they can converge close to the minimum norm solution. This is different from the behavior of gradient descent, which converges to the minimum norm solution only when started at a point in the range space of the feature matrix, and is thus more sensitive to initialization.
NeurIPS 2020
References in corpus (25)
- Instance Normalization: The Missing Ingredient for Fast Stylization
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- Neural Tangent Kernel: Convergence and Generalization in Neural Networks
- Understanding deep learning requires rethinking generalization
- Benign Overfitting in Linear Regression
- Gradient Descent Provably Optimizes Over-parameterized Neural Networks
- Dropout Training as Adaptive Regularization
- No Spurious Local Minima in Nonconvex Low Rank Problems: A Unified Geometric Analysis
- Understanding and Improving Layer Normalization
- AdaGrad stepsizes: Sharp convergence over nonconvex landscapes
- Characterizing Implicit Bias in Terms of Optimization Geometry
- Gradient Descent Converges to Minimizers
- Norm matters: efficient and accurate normalization schemes in deep networks
- Theoretical Analysis of Auto Rate-Tuning by Batch Normalization
- Implicit Regularization in Deep Matrix Factorization
- The Phase Transition of Matrix Recovery from Gaussian Measurements Matches the Minimax MSE of Matrix Denoising
- WNGrad: Learn the Learning Rate in Gradient Descent
- Towards Understanding Regularization in Batch Normalization
- On the Implicit Bias of Dropout
- Theoretical Issues in Deep Networks: Approximation, Optimization and Generalization
- Luck Matters: Understanding Training Dynamics of Deep ReLU Networks
- A Continuous-Time View of Early Stopping for Least Squares
- Optimization Theory for ReLU Neural Networks Trained with Normalization Layers
- Dropout: Explicit Forms and Capacity Control
- Student Specialization in Deep ReLU Networks With Finite Width and Input Dimension