Identity Crisis: Memorization and Generalization under Extreme Overparameterization
arXiv:1902.04698
Abstract
We study the interplay between memorization and generalization of overparameterized networks in the extreme case of a single training example and an identity-mapping task. We examine fully-connected and convolutional networks (FCN and CNN), both linear and nonlinear, initialized randomly and then trained to minimize the reconstruction error. The trained networks stereotypically take one of two forms: the constant function (memorization) and the identity function (generalization). We formally characterize generalization in single-layer FCNs and CNNs. We show empirically that different architectures exhibit strikingly different inductive biases. For example, CNNs of up to 10 layers are able to generalize from a single example, whereas FCNs cannot learn the identity function reliably from 60k examples. Deeper CNNs often fail, but nonetheless do astonishing work to memorize the training output: because CNN biases are location invariant, the model must progressively grow an output pattern from the image boundaries via the coordination of many layers. Our work helps to quantify and visualize the sensitivity of inductive biases to architectural choices such as depth, kernel width, and number of channels.
ICLR 2020
References in corpus (15)
- Conditional Generative Adversarial Nets
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- A Convergence Theory for Deep Learning via Over-Parameterization
- A Closer Look at Memorization in Deep Networks
- A PAC-Bayesian Approach to Spectrally-Normalized Margin Bounds for Neural Networks
- Learning Overparameterized Neural Networks via Stochastic Gradient Descent on Structured Data
- Computing Nonvacuous Generalization Bounds for Deep (Stochastic) Neural Networks with Many More Parameters than Training Data
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Learning and Generalization in Overparameterized Neural Networks, Going Beyond Two Layers
- Fisher-Rao Metric, Geometry, and Complexity of Neural Networks
- Non-Vacuous Generalization Bounds at the ImageNet Scale: A PAC-Bayesian Compression Approach
- On exponential convergence of SGD in non-convex over-parametrized learning
- Overparameterized Nonlinear Learning: Gradient Descent Takes the Shortest Path?
- Memorization in Overparameterized Autoencoders
- Minimum weight norm models do not always generalize well for over-parameterized problems
Cited by in corpus (14)
- The Modern Mathematics of Deep Learning
- Deep Gamblers: Learning to Abstain with Portfolio Theory
- Memorization in Overparameterized Autoencoders
- What they do when in doubt: a study of inductive biases in seq2seq learners
- Luck Matters: Understanding Training Dynamics of Deep ReLU Networks
- A Unifying View on Implicit Bias in Training Linear Neural Networks
- Inductive Bias of Multi-Channel Linear Convolutional Networks with Bounded Weight Norm
- Neural Networks Trained on Natural Scenes Exhibit Gestalt Closure
- Which Minimizer Does My Neural Network Converge To?
- Rethink the Connections among Generalization, Memorization and the Spectral Bias of DNNs
- Rate-Regularization and Generalization in VAEs
- A Primer for Neural Arithmetic Logic Modules
- Emergent Properties of Finetuned Language Representation Models
- On Alignment in Deep Linear Neural Networks