Generalization in Deep Learning
arXiv:1710.05468 · doi:10.1017/9781009025096.003
Abstract
This paper provides theoretical insights into why and how deep learning can generalize well, despite its large capacity, complexity, possible algorithmic instability, nonrobustness, and sharp minima, responding to an open question in the literature. We also discuss approaches to provide non-vacuous generalization guarantees for deep learning. Based on theoretical observations, we propose new open problems and discuss the limitations of our results.
Published by Cambridge University Press. BibTeX of this paper is available at: https://people.csail.mit.edu/kawaguch/bibtex.html
References in corpus (10)
- The Loss Surfaces of Multilayer Networks
- Train longer, generalize better: closing the generalization gap in large batch training of neural networks
- A Closer Look at Memorization in Deep Networks
- Path-SGD: Path-Normalized Optimization in Deep Neural Networks
- Towards Understanding Generalization of Deep Learning: Perspective of Loss Landscapes
- APAC: Augmented PAttern Classification with Neural Networks
- On the Computational Efficiency of Training Neural Networks
- Norm-Based Capacity Control in Neural Networks
- Fast Rates for Empirical Risk Minimization of Strict Saddle Problems
- Generalization in Machine Learning via Analytical Learning Theory
Cited by in corpus (17)
- Learning ReLU Networks on Linearly Separable Data: Algorithm, Optimality, and Generalization
- Unsupervised phase discovery with deep anomaly detection
- Deep learning of quantum entanglement from incomplete measurements
- ExpandNets: Linear Over-parameterization to Train Compact Convolutional Networks
- Topology of Learning in Artificial Neural Networks
- Hutchinson Trace Estimation for High-Dimensional and High-Order Physics-Informed Neural Networks
- The Future of Intelligent Healthcare: A Systematic Analysis and Discussion on the Integration and Impact of Robots Using Large Language Models for Healthcare
- Experiment data-driven modeling of tokamak discharge in EAST
- Online Deep Learning from Doubly-Streaming Data
- Statistical Complexity of Quantum Learning
- Gradient-Free Learning Based on the Kernel and the Range Space
- Analytic Network Learning
- Gabor frames and deep scattering networks in audio processing
- Do Sharpness-based Optimizers Improve Generalization in Medical Image Analysis?
- On Neural Inertial Classification Networks for Pedestrian Activity Recognition
- The Renyi Gaussian Process: Towards Improved Generalization
- Should All Noises Be Treated Equally: Impact of Input Noise Variability on Neural Network Robustness