Towards an Understanding of Residual Networks Using Neural Tangent Hierarchy (NTH)
arXiv:2007.03714 · doi:10.4208/csiam-am.SO-2021-0053
Abstract
Gradient descent yields zero training loss in polynomial time for deep neural networks despite non-convex nature of the objective function. The behavior of network in the infinite width limit trained by gradient descent can be described by the Neural Tangent Kernel (NTK) introduced in \cite{Jacot2018Neural}. In this paper, we study dynamics of the NTK for finite width Deep Residual Network (ResNet) using the neural tangent hierarchy (NTH) proposed in \cite{Huang2019Dynamics}. For a ResNet with smooth and Lipschitz activation function, we reduce the requirement on the layer width with respect to the number of training samples from quartic to cubic. Our analysis suggests strongly that the particular skip-connection structure of ResNet is the main reason for its triumph over fully-connected network.
72 pages, 1 figure
References in corpus (12)
- On Exact Computation with an Infinitely Wide Neural Net
- Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient Descent
- How to Escape Saddle Points Efficiently
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Escaping From Saddle Points --- Online Stochastic Gradient for Tensor Decomposition
- Scaling Limits of Wide Neural Networks with Weight Sharing: Gaussian Process Behavior, Gradient Independence, and Neural Tangent Kernel Derivation
- Learning One-hidden-layer Neural Networks with Landscape Design
- A Comparative Analysis of the Optimization and Generalization Property of Two-layer Neural Network and Random Feature Models Under Gradient Descent Dynamics
- Fast Convergence of Natural Gradient Descent for Overparameterized Neural Networks
- Quadratic Suffices for Over-parametrization via Matrix Chernoff Bound
- A Priori Estimates of the Population Risk for Residual Networks
- Critical Points of Neural Networks: Analytical Forms and Landscape Properties