Disentangling Trainability and Generalization in Deep Neural Networks
arXiv:1912.13053
Abstract
A longstanding goal in the theory of deep learning is to characterize the conditions under which a given neural network architecture will be trainable, and if so, how well it might generalize to unseen data. In this work, we provide such a characterization in the limit of very wide and very deep networks, for which the analysis simplifies considerably. For wide networks, the trajectory under gradient descent is governed by the Neural Tangent Kernel (NTK), and for deep networks the NTK itself maintains only weak data dependence. By analyzing the spectrum of the NTK, we formulate necessary conditions for trainability and generalization across a range of architectures, including Fully Connected Networks (FCNs) and Convolutional Neural Networks (CNNs). We identify large regions of hyperparameter space for which networks can memorize the training set but completely fail to generalize. We find that CNNs without global average pooling behave almost identically to FCNs, but that CNNs with pooling have markedly different and often better generalization performance. These theoretical results are corroborated experimentally on CIFAR10 for a variety of network architectures and we include a colab notebook that reproduces the essential results of the paper.
22 pages, 3 figures, ICML 2020. Associated Colab notebook at https://colab.research.google.com/github/google/neural-tangents/blob/master/notebooks/Disentangling_Trainability_and_Generalization.ipynb
Cited by in corpus (9)
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations
- GradSign: Model Performance Inference with Theoretical Insights
- Increasing Depth Leads to U-Shaped Test Risk in Over-parameterized Convolutional Networks
- The Care Label Concept: A Certification Suite for Trustworthy and Resource-Aware Machine Learning
- Gradients are Not All You Need
- Large-Dimensional Random Matrix Theory and Its Applications in Deep Learning and Wireless Communications
- Adaptive Latent Space Tuning for Non-Stationary Distributions
- Rapid training of deep neural networks without skip connections or normalization layers using Deep Kernel Shaping
- Disentangling deep neural networks with rectified linear units using duality