Generalization in Deep Networks: The Role of Distance from Initialization
arXiv:1901.01672
Abstract
Why does training deep neural networks using stochastic gradient descent (SGD) result in a generalization error that does not worsen with the number of parameters in the network? To answer this question, we advocate a notion of effective model capacity that is dependent on {\em a given random initialization of the network} and not just the training algorithm and the data distribution. We provide empirical evidences that demonstrate that the model capacity of SGD-trained deep networks is in fact restricted through implicit regularization of {\em the distance from the initialization}. We also provide theoretical arguments that further highlight the need for initialization-dependent notions of model capacity. We leave as open questions how and why distance from initialization is regularized, and whether it is sufficient to explain generalization.
Spotlight paper at NeurIPS 2017 workshop on Deep Learning: Bridging Theory and Practice
References in corpus (1)
Cited by in corpus (12)
- Fantastic Generalization Measures and Where to Find Them
- Deployment-Efficient Reinforcement Learning via Model-Based Offline Optimization
- Observational Overfitting in Reinforcement Learning
- The intriguing role of module criticality in the generalization of deep networks
- AdjointNet: Constraining machine learning models with physics-based codes
- Robustness to Pruning Predicts Generalization in Deep Neural Networks
- A Theoretical Analysis of Fine-tuning with Linear Teachers
- Measuring Generalization with Optimal Transport
- Improved generalization by noise enhancement
- Practical Assessment of Generalization Performance Robustness for Deep Networks via Contrastive Examples
- Explaining generalization in deep learning: progress and fundamental limits
- Understanding the Role of Adversarial Regularization in Supervised Learning