Understanding deep learning requires rethinking generalization
arXiv:1611.03530
Abstract
Despite their massive size, successful deep artificial neural networks can exhibit a remarkably small difference between training and test performance. Conventional wisdom attributes small generalization error either to properties of the model family, or to the regularization techniques used during training. Through extensive systematic experiments, we show how these traditional approaches fail to explain why large neural networks generalize well in practice. Specifically, our experiments establish that state-of-the-art convolutional networks for image classification trained with stochastic gradient methods easily fit a random labeling of the training data. This phenomenon is qualitatively unaffected by explicit regularization, and occurs even if we replace the true images by completely unstructured random noise. We corroborate these experimental findings with a theoretical construction showing that simple depth two neural networks already have perfect finite sample expressivity as soon as the number of parameters exceeds the number of data points as it usually does in practice. We interpret our experimental findings by comparison with traditional models.
Published in ICLR 2017
References in corpus (2)
Cited by in corpus (52)
- Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey
- Transferable Clean-Label Poisoning Attacks on Deep Neural Nets
- Overfitting Mechanism and Avoidance in Deep Neural Networks
- On the Origin of Deep Learning
- Globally Optimal Gradient Descent for a ConvNet with Gaussian Inputs
- Analysis and Optimization of Convolutional Neural Network Architectures
- Generalization Bounds of SGLD for Non-convex Learning: Two Theoretical Viewpoints
- An Improved Analysis of Training Over-parameterized Deep Neural Networks
- Superposition of many models into one
- A mean-field limit for certain deep neural networks
- SGD Learns Over-parameterized Networks that Provably Generalize on Linearly Separable Data
- On the Robustness of Convolutional Neural Networks to Internal Architecture and Weight Perturbations
- Explicitizing an Implicit Bias of the Frequency Principle in Two-layer Neural Networks
- Classification regions of deep neural networks
- Theoretical properties of the global optimizer of two layer neural network
- Deep Neural Networks for Marine Debris Detection in Sonar Images
- Porcupine Neural Networks: (Almost) All Local Optima are Global
- Negative eigenvalues of the Hessian in deep neural networks
- Intriguing Properties of Adversarial Examples
- Stability and Generalization of Learning Algorithms that Converge to Global Optima
- Weight-space symmetry in deep networks gives rise to permutation saddles, connected by equal-loss valleys across the loss landscape
- How do infinite width bounded norm networks look in function space?
- Segmentation-Aware Image Denoising without Knowing True Segmentation
- Stability and Generalization of Graph Convolutional Neural Networks
- Neural Entropic Estimation: A faster path to mutual information estimation
- Positively Scale-Invariant Flatness of ReLU Neural Networks
- Understanding Adversarial Robustness Through Loss Landscape Geometries
- On Dropout and Nuclear Norm Regularization
- Quasi-potential as an implicit regularizer for the loss function in the stochastic gradient descent
- Data-Efficient Mutual Information Neural Estimator
- On Scalable and Efficient Computation of Large Scale Optimal Transport
- Universal Supervised Learning for Individual Data
- Label-Noise Robust Multi-Domain Image-to-Image Translation
- Deep learning in bioinformatics: introduction, application, and perspective in big data era
- Ensemble Model Patching: A Parameter-Efficient Variational Bayesian Neural Network
- Deep learning methods based on cross-section images for predicting effective thermal conductivity of composites
- Distributed Optimization for Over-Parameterized Learning
- Rethinking the Artificial Neural Networks: A Mesh of Subnets with a Central Mechanism for Storing and Predicting the Data
- From Data Quality to Model Quality: an Exploratory Study on Deep Learning
- The Geometry of Deep Networks: Power Diagram Subdivision
- GM-Net: Learning Features with More Efficiency
- Deep Learning: Generalization Requires Deep Compositional Feature Space Design
- An Essay on Optimization Mystery of Deep Learning
- Improved visible to IR image transformation using synthetic data augmentation with cycle-consistent adversarial networks
- The role of a layer in deep neural networks: a Gaussian Process perspective
- BPGrad: Towards Global Optimality in Deep Learning via Branch and Pruning
- On the potential for open-endedness in neural networks
- Understanding the Behaviour of the Empirical Cross-Entropy Beyond the Training Distribution
- Deep Optimization for Spectrum Repacking
- Computer activity learning from system call time series
- Effect of Various Regularizers on Model Complexities of Neural Networks in Presence of Input Noise
- Nonparametric Online Learning Using Lipschitz Regularized Deep Neural Networks