Dynamics of Deep Neural Networks and Neural Tangent Hierarchy
arXiv:1909.08156
Abstract
The evolution of a deep neural network trained by the gradient descent can be described by its neural tangent kernel (NTK) as introduced in [20], where it was proven that in the infinite width limit the NTK converges to an explicit limiting kernel and it stays constant during training. The NTK was also implicit in some other recent papers [6,13,14]. In the overparametrization regime, a fully-trained deep neural network is indeed equivalent to the kernel regression predictor using the limiting NTK. And the gradient descent achieves zero training loss for a deep overparameterized neural network. However, it was observed in [5] that there is a performance gap between the kernel regression using the limiting NTK and the deep neural networks. This performance gap is likely to originate from the change of the NTK along training due to the finite width effect. The change of the NTK along the training is central to describe the generalization features of deep neural networks. In the current paper, we study the dynamic of the NTK for finite width deep fully-connected neural networks. We derive an infinite hierarchy of ordinary differential equations, the neural tangent hierarchy (NTH) which captures the gradient descent dynamic of the deep neural network. Moreover, under certain conditions on the neural network width and the data set dimension, we prove that the truncated hierarchy of NTH approximates the dynamic of the NTK up to arbitrary precision. This description makes it possible to directly study the change of the NTK for deep neural networks, and sheds light on the observation that deep neural networks outperform kernel regressions using the corresponding limiting NTK.
References in corpus (9)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- The Loss Surfaces of Multilayer Networks
- Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
- Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
- A mean-field limit for certain deep neural networks
- Mean Field Limit of the Learning Dynamics of Multilayer Neural Networks
- Gradient Descent Finds Global Minima for Generalizable Deep Neural Networks of Practical Sizes
Cited by in corpus (22)
- The large learning rate phase of deep learning: the catapult mechanism
- Tensor Programs II: Neural Tangent Kernel for Any Architecture
- Asymptotics of Wide Networks from Feynman Diagrams
- Gradient Starvation: A Learning Proclivity in Neural Networks
- Frequency Bias in Neural Networks for Input of Non-Uniform Density
- Bayesian Deep Ensembles via the Neural Tangent Kernel
- Feature Learning in Infinite-Width Neural Networks
- Spectra of the Conjugate Kernel and Neural Tangent Kernel for linear-width neural networks
- Asymptotics of Wide Convolutional Neural Networks
- On the Neural Tangent Kernel of Deep Networks with Orthogonal Initialization
- Generalization bounds for deep learning
- Towards Understanding Hierarchical Learning: Benefits of Neural Representations
- Towards Deepening Graph Neural Networks: A GNTK-based Optimization Perspective
- The asymptotic spectrum of the Hessian of DNN throughout training
- Implicit bias of deep linear networks in the large learning rate phase
- DNN-Based Topology Optimisation: Spatial Invariance and Neural Tangent Kernel
- On the Empirical Neural Tangent Kernel of Standard Finite-Width Convolutional Neural Network Architectures
- Dynamically Stable Infinite-Width Limits of Neural Classifiers
- Label-Aware Neural Tangent Kernel: Toward Better Generalization and Local Elasticity
- Predicting the outputs of finite deep neural networks trained with noisy gradients
- Notes on Deep Learning Theory
- Whitening and second order optimization both make information in the dataset unusable during training, and can reduce or prevent generalization