Continuous vs. Discrete Optimization of Deep Neural Networks
arXiv:2107.06608
Abstract
Existing analyses of optimization in deep learning are either continuous, focusing on (variants of) gradient flow, or discrete, directly treating (variants of) gradient descent. Gradient flow is amenable to theoretical analysis, but is stylized and disregards computational efficiency. The extent to which it represents gradient descent is an open question in the theory of deep learning. The current paper studies this question. Viewing gradient descent as an approximate numerical solution to the initial value problem of gradient flow, we find that the degree of approximation depends on the curvature around the gradient flow trajectory. We then show that over deep neural networks with homogeneous activations, gradient flow trajectories enjoy favorable curvature, suggesting they are well approximated by gradient descent. This finding allows us to translate an analysis of gradient flow over deep linear neural networks into a guarantee that gradient descent efficiently converges to global minimum almost surely under random initialization. Experiments suggest that over simple deep neural networks, gradient descent with conventional step size is indeed close to gradient flow. We hypothesize that the theory of gradient flows will unravel mysteries behind deep learning.
Published as spotlight paper at the conference on Neural Information Processing Systems (NeurIPS) 2021
References in corpus (19)
- A Convergence Theory for Deep Learning via Over-Parameterization
- Gradient descent aligns the layers of deep linear networks
- Gradient Descent Maximizes the Margin of Homogeneous Neural Networks
- Understanding the Acceleration Phenomenon via High-Resolution Differential Equations
- Implicit Regularization in Deep Learning May Not Be Explainable by Norms
- The large learning rate phase of deep learning: the catapult mechanism
- An analytic theory of generalization dynamics and transfer learning in deep linear networks
- On the Origin of Implicit Regularization in Stochastic Gradient Descent
- Provable Benefit of Orthogonal Initialization in Optimizing Deep Linear Networks
- The Break-Even Point on Optimization Trajectories of Deep Neural Networks
- Width Provably Matters in Optimization for Deep Linear Neural Networks
- Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability
- Analysis of the Gradient Descent Algorithm for a Deep Neural Network Model with Skip-connections
- Integration Methods and Accelerated Optimization Algorithms
- Shadowing Properties of Optimization Algorithms
- On the Global Convergence of Training Deep Linear ResNets
- Implicit Regularization in ReLU Networks with the Square Loss
- Global Convergence of Gradient Descent for Deep Linear Residual Networks
- A Convergence Theory Towards Practical Over-parameterized Deep Neural Networks