Shape Matters: Understanding the Implicit Bias of the Noise Covariance
arXiv:2006.08680
Abstract
The noise in stochastic gradient descent (SGD) provides a crucial implicit regularization effect for training overparameterized models. Prior theoretical work largely focuses on spherical Gaussian noise, whereas empirical studies demonstrate the phenomenon that parameter-dependent noise -- induced by mini-batches or label perturbation -- is far more effective than Gaussian noise. This paper theoretically characterizes this phenomenon on a quadratically-parameterized model introduced by Vaskevicius et el. and Woodworth et el. We show that in an over-parameterized setting, SGD with label noise recovers the sparse ground-truth with an arbitrary initialization, whereas SGD with Gaussian noise or gradient descent overfits to dense solutions with large norms. Our analysis reveals that parameter-dependent noise introduces a bias towards local minima with smaller noise variance, whereas spherical Gaussian noise does not. Code for our project is publicly available.
References in corpus (12)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Understanding deep learning requires rethinking generalization
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
- Path-SGD: Path-Normalized Optimization in Deep Neural Networks
- Generalization Bounds of SGLD for Non-convex Learning: Two Theoretical Viewpoints
- Theory of Deep Learning III: explaining the non-overfitting puzzle
- Information-Theoretic Generalization Bounds for SGLD via Data-Dependent Estimates
- Lexicographic and Depth-Sensitive Margins in Homogeneous and Non-Homogeneous Deep Models
- How noise affects the Hessian spectrum in overparameterized neural networks
- On Dropout and Nuclear Norm Regularization
- Disentangling Adaptive Gradient Methods from Learning Rates
Cited by in corpus (13)
- On the Opportunities and Risks of Foundation Models
- Self-supervised Learning is More Robust to Dataset Imbalance
- Stochasticity helps to navigate rough landscapes: comparing gradient-descent-based algorithms in the phase retrieval problem
- Strength of Minibatch Noise in SGD
- Label Noise SGD Provably Prefers Flat Global Minimizers
- Stochastic Training is Not Necessary for Generalization
- Learning Over-Parametrized Two-Layer ReLU Neural Networks beyond NTK
- What Happens after SGD Reaches Zero Loss? --A Mathematical Framework
- Implicit Bias of SGD for Diagonal Linear Networks: a Provable Benefit of Stochasticity
- Optimal Gradient-based Algorithms for Non-concave Bandit Optimization
- Noisy Gradient Descent Converges to Flat Minima for Nonconvex Matrix Factorization
- Going Beyond Linear RL: Sample Efficient Neural Function Approximation
- Towards Understanding Generalization via Decomposing Excess Risk Dynamics