Gradient Descent with Early Stopping is Provably Robust to Label Noise for Overparameterized Neural Networks
arXiv:1903.11680
Abstract
Modern neural networks are typically trained in an over-parameterized regime where the parameters of the model far exceed the size of the training data. Such neural networks in principle have the capacity to (over)fit any set of labels including pure noise. Despite this, somewhat paradoxically, neural network models trained via first-order methods continue to predict well on yet unseen test data. This paper takes a step towards demystifying this phenomena. Under a rich dataset model, we show that gradient descent is provably robust to noise/corruption on a constant fraction of the labels despite overparameterization. In particular, we prove that: (i) In the first few iterations where the updates are still in the vicinity of the initialization gradient descent only fits to the correct labels essentially ignoring the noisy labels. (ii) to start to overfit to the noisy labels network must stray rather far from from the initialization which can only occur after many more iterations. Together, these results show that gradient descent with early stopping is provably robust to label noise and shed light on the empirical robustness of deep networks as well as commonly adopted heuristics to prevent overfitting.
References in corpus (5)
- Training Deep Neural Networks on Noisy Labels with Bootstrapping
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
- Robust Regression via Hard Thresholding
- Generalization Guarantees for Neural Networks via Harnessing the Low-rank Structure of the Jacobian
Cited by in corpus (21)
- Denoising and Regularization via Exploiting the Structural Bias of Convolutional Generators
- Generalization Guarantees for Neural Networks via Harnessing the Low-rank Structure of the Jacobian
- Distillation Early Stopping? Harvesting Dark Knowledge Utilizing Anisotropic Information Retrieval For Overparameterized Neural Network
- Compressive sensing with un-trained neural networks: Gradient descent finds the smoothest approximation
- Learning Not to Learn in the Presence of Noisy Labels
- How benign is benign overfitting?
- Rethinking Influence Functions of Neural Networks in the Over-parameterized Regime
- Coresets for Robust Training of Neural Networks against Noisy Labels
- On the Role of Dataset Quality and Heterogeneity in Model Confidence
- Adaptive Precision Training (AdaPT): A dynamic fixed point quantized training approach for DNNs
- Exploring Weight Importance and Hessian Bias in Model Pruning
- Deep Transfer Learning for Automated Diagnosis of Skin Lesions from Photographs
- Image recognition from raw labels collected without annotators
- KNN-enhanced Deep Learning Against Noisy Labels
- Learning with Noisy Labels by Efficient Transition Matrix Estimation to Combat Label Miscorrection
- Heavy-tailed Streaming Statistical Estimation
- Phases of learning dynamics in artificial neural networks: with or without mislabeled data
- A Tale Of Two Long Tails
- Robustness and Reliability When Training With Noisy Labels
- Constrained Instance and Class Reweighting for Robust Learning under Label Noise
- Generalization by Recognizing Confusion