On the Optimization of Deep Networks: Implicit Acceleration by Overparameterization
arXiv:1802.06509
Abstract
Conventional wisdom in deep learning states that increasing depth improves expressiveness but complicates optimization. This paper suggests that, sometimes, increasing depth can speed up optimization. The effect of depth on optimization is decoupled from expressiveness by focusing on settings where additional layers amount to overparameterization - linear neural networks, a well-studied model. Theoretical analysis, as well as experiments, show that here depth acts as a preconditioner which may accelerate convergence. Even on simple convex problems such as linear regression with loss, , gradient descent can benefit from transitioning to a non-convex overparameterized objective, more than it would from some common acceleration schemes. We also prove that it is mathematically impossible to obtain the acceleration effect of overparametrization via gradients of any regularizer.
Published at the International Conference on Machine Learning (ICML) 2018
References in corpus (5)
Cited by in corpus (62)
- Knowledge Distillation: A Survey
- Beyond English-Centric Multilingual Machine Translation
- Gradient Descent Provably Optimizes Over-parameterized Neural Networks
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Gradient Descent with Early Stopping is Provably Robust to Label Noise for Overparameterized Neural Networks
- ResNet with one-neuron hidden layers is a Universal Approximator
- Gradient Descent Happens in a Tiny Subspace
- Deep matrix factorizations
- Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
- A Modern Take on the Bias-Variance Tradeoff in Neural Networks
- Overparameterized Neural Networks Implement Associative Memory
- The learnability scaling of quantum states: restricted Boltzmann machines
- Stiffness: A New Perspective on Generalization in Neural Networks
- Effect of Depth and Width on Local Minima in Deep Learning
- Width Provably Matters in Optimization for Deep Linear Neural Networks
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient Descent
- Distillation Early Stopping? Harvesting Dark Knowledge Utilizing Anisotropic Information Retrieval For Overparameterized Neural Network
- Exponential Convergence Time of Gradient Descent for One-Dimensional Deep Linear Neural Networks
- Algorithmic Regularization in Learning Deep Homogeneous Models: Layers are Automatically Balanced
- Recurrent Quantum Neural Networks
- Network Pruning That Matters: A Case Study on Retraining Variants
- Lookahead: A Far-Sighted Alternative of Magnitude-based Pruning
- Integrating Deep Neural Networks with Full-waveform Inversion: Reparametrization, Regularization, and Uncertainty Quantification
- The Benefits of Over-parameterization at Initialization in Deep ReLU Networks
- Revealing the Structure of Deep Neural Networks via Convex Duality
- Neural Mechanics: Symmetry and Broken Conservation Laws in Deep Learning Dynamics
- The Implicit Bias of Depth: How Incremental Learning Drives Generalization
- An Empirical Study of Large-Batch Stochastic Gradient Descent with Structured Covariance Noise
- Revisiting Landscape Analysis in Deep Neural Networks: Eliminating Decreasing Paths to Infinity
- Depth creates no more spurious local minima
- Neural Empirical Bayes
- Deep Neural Networks with Multi-Branch Architectures Are Less Non-Convex
- Interpreting Deep Learning: The Machine Learning Rorschach Test?
- A Bayesian Perspective of Convolutional Neural Networks through a Deconvolutional Generative Model
- Finding the Needle in the Haystack with Convolutions: on the benefits of architectural bias
- Auxiliary Learning for Deep Multi-task Learning
- Hessian based analysis of SGD for Deep Nets: Dynamics and Generalization
- Implicit Under-Parameterization Inhibits Data-Efficient Deep Reinforcement Learning
- Noether: The More Things Change, the More Stay the Same
- A Differential Topological View of Challenges in Learning with Feedforward Neural Networks
- Spectral Geometric Matrix Completion
- Distributed Optimization for Over-Parameterized Learning
- Understanding over-parameterized deep networks by geometrization
- Thompson Sampling for Noncompliant Bandits
- Active multi-fidelity Bayesian online changepoint detection
- Local Methods with Adaptivity via Scaling
- Noether's Learning Dynamics: Role of Symmetry Breaking in Neural Networks
- A Convergence Theory Towards Practical Over-parameterized Deep Neural Networks
- Two-Level K-FAC Preconditioning for Deep Learning
- A Modular Analysis of Provable Acceleration via Polyak's Momentum: Training a Wide ReLU Network and a Deep Linear Network
- Volumization as a Natural Generalization of Weight Decay
- Painless step size adaptation for SGD
- Training Linear Neural Networks: Non-Local Convergence and Complexity Results
- Directional Convergence Analysis under Spherically Symmetric Distribution
- Towards Better Generalization: BP-SVRG in Training Deep Neural Networks
- Associated Learning: Decomposing End-to-end Backpropagation based on Auto-encoders and Target Propagation
- Understanding Modern Techniques in Optimization: Frank-Wolfe, Nesterov's Momentum, and Polyak's Momentum
- Implicit Acceleration and Feature Learning in Infinitely Wide Neural Networks with Bottlenecks
- Layer Dynamics of Linearised Neural Nets
- Deep Learning: a new definition of artificial neuron with double weight