Geometry of Optimization and Implicit Regularization in Deep Learning
arXiv:1705.03071
Abstract
We argue that the optimization plays a crucial role in generalization of deep learning models through implicit regularization. We do this by demonstrating that generalization ability is not controlled by network size but rather by some other implicit control. We then demonstrate how changing the empirical optimization procedure can improve generalization, even if actual optimization quality is not affected. We do so by studying the geometry of the parameter space of deep networks, and devising an optimization algorithm attuned to this geometry.
This survey chapter was done as a part of Intel Collaborative Research institute for Computational Intelligence (ICRI-CI) "Why & When Deep Learning works -- looking inside Deep Learning" compendium with the generous support of ICRI-CI. arXiv admin note: substantial text overlap with arXiv:1506.02617
References in corpus (1)
Cited by in corpus (43)
- Unsupervised Domain Adaptation through Self-Supervision
- Characterizing Implicit Bias in Terms of Optimization Geometry
- Scaling description of generalization with number of parameters in deep learning
- From Variational to Deterministic Autoencoders
- Double Trouble in Double Descent : Bias and Variance(s) in the Lazy Regime
- Deep Learning Theory Review: An Optimal Control and Dynamical Systems Perspective
- Gradient Starvation: A Learning Proclivity in Neural Networks
- Approximation and Estimation for High-Dimensional Deep Learning Networks
- What Kinds of Functions do Deep Neural Networks Learn? Insights from Variational Spline Theory
- Implicit Bias of Gradient Descent on Linear Convolutional Networks
- Bad Global Minima Exist and SGD Can Reach Them
- Stochastic Mirror Descent on Overparameterized Nonlinear Models: Convergence, Implicit Regularization, and Generalization
- Lexicographic and Depth-Sensitive Margins in Homogeneous and Non-Homogeneous Deep Models
- Implicit Regularization of Discrete Gradient Dynamics in Linear Neural Networks
- NeurIPS 2020 Competition: Predicting Generalization in Deep Learning
- Structure Probing Neural Network Deflation
- Mean-Field Neural ODEs via Relaxed Optimal Control
- On Convergence and Generalization of Dropout Training
- An Empirical Study of Large-Batch Stochastic Gradient Descent with Structured Covariance Noise
- Reproducing Activation Function for Deep Learning
- The Local Elasticity of Neural Networks
- Quasi-potential as an implicit regularizer for the loss function in the stochastic gradient descent
- On implicit regularization: Morse functions and applications to matrix factorization
- PAC-Bayesian Margin Bounds for Convolutional Neural Networks
- Can Implicit Bias Explain Generalization? Stochastic Convex Optimization as a Case Study
- Implicit Regularization via Neural Feature Alignment
- The Discovery of Dynamics via Linear Multistep Methods and Deep Learning: Error Estimation
- Learning Dynamics of Linear Denoising Autoencoders
- On the Double Descent of Random Features Models Trained with SGD
- Perspective: A Phase Diagram for Deep Learning unifying Jamming, Feature Learning and Lazy Training
- Implicit Sparse Regularization: The Impact of Depth and Early Stopping
- Iterative regularization for convex regularizers
- The Implicit Bias of AdaGrad on Separable Data
- Implicit Regularization and Entrywise Convergence of Riemannian Optimization for Low Tucker-Rank Tensor Completion
- Robust Implicit Regularization via Weight Normalization
- Residual Networks: Lyapunov Stability and Convex Decomposition
- Effective Regularization Through Loss-Function Metalearning
- Minimum complexity interpolation in random features models
- Analytic expressions for the output evolution of a deep neural network
- How isotropic kernels perform on simple invariants
- Understanding the Behaviour of the Empirical Cross-Entropy Beyond the Training Distribution
- Step Size Matters in Deep Learning
- Comments on Leo Breiman's paper 'Statistical Modeling: The Two Cultures' (Statistical Science, 2001, 16(3), 199-231)