An analytic theory of generalization dynamics and transfer learning in deep linear networks
arXiv:1809.10374
Abstract
Much attention has been devoted recently to the generalization puzzle in deep learning: large, deep networks can generalize well, but existing theories bounding generalization error are exceedingly loose, and thus cannot explain this striking performance. Furthermore, a major hope is that knowledge may transfer across tasks, so that multi-task learning can improve generalization on individual tasks. However we lack analytic theories that can quantitatively predict how the degree of knowledge transfer depends on the relationship between the tasks. We develop an analytic theory of the nonlinear dynamics of generalization in deep linear networks, both within and across tasks. In particular, our theory provides analytic solutions to the training and testing error of deep networks as a function of training time, number of examples, network size and initialization, and the task structure and SNR. Our theory reveals that deep networks progressively learn the most important task structure first, so that generalization error at the early stopping time primarily depends on task structure and is independent of network size. This suggests any tight bound on generalization error must take into account task structure, and explains observations about real data being learned faster than random data. Intriguingly our theory also reveals the existence of a learning algorithm that proveably out-performs neural network training through gradient descent. Finally, for transfer learning, our theory reveals that knowledge transfer depends sensitively, but computably, on the SNRs and input feature alignments of pairs of tasks.
ICLR 2019, 20 pages
References in corpus (2)
Cited by in corpus (35)
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Intrinsic dimension of data representations in deep neural networks
- Memorizing without overfitting: Bias, variance, and interpolation in over-parameterized models
- Implicit Regularization in Deep Matrix Factorization
- Implicit Regularization in Deep Learning May Not Be Explainable by Norms
- Environmental drivers of systematicity and generalization in a situated agent
- What shapes feature representations? Exploring datasets, architectures, and training
- Optimal Regularization Can Mitigate Double Descent
- Gradient Starvation: A Learning Proclivity in Neural Networks
- Understanding self-supervised Learning Dynamics without Contrastive Pairs
- Understanding Self-supervised Learning with Dual Deep Networks
- Triple descent and the two kinds of overfitting: Where & why do they appear?
- Implicit Regularization of Discrete Gradient Dynamics in Linear Neural Networks
- Probing transfer learning with a model of synthetic correlated datasets
- Technical Considerations for Semantic Segmentation in MRI using Convolutional Neural Networks
- A Random Matrix Perspective on Mixtures of Nonlinearities for Deep Learning
- Phase Transitions in Transfer Learning for High-Dimensional Perceptrons
- The Surprising Simplicity of the Early-Time Learning Dynamics of Neural Networks
- On the geometry of generalization and memorization in deep neural networks
- PolyGAN: High-Order Polynomial Generators
- Generalization in multitask deep neural classifiers: a statistical physics approach
- Implicit Regularization via Neural Feature Alignment
- A Sample Complexity Separation between Non-Convex and Convex Meta-Learning
- Continuous vs. Discrete Optimization of Deep Neural Networks
- Towards Demystifying Representation Learning with Non-contrastive Self-supervision
- Abstraction Mechanisms Predict Generalization in Deep Neural Networks
- Trap of Feature Diversity in the Learning of MLPs
- An analytic theory of shallow networks dynamics for hinge loss classification
- The Geometry of Over-parameterized Regression and Adversarial Perturbations
- Implicit Regularization in Tensor Factorization
- An Empirical Study on the Relation between Network Interpretability and Adversarial Robustness
- Student Specialization in Deep ReLU Networks With Finite Width and Input Dimension
- Zero-shot task adaptation by homoiconic meta-mapping
- Post-Workshop Report on Science meets Engineering in Deep Learning, NeurIPS 2019, Vancouver
- Geometry Perspective Of Estimating Learning Capability Of Neural Networks