Simple and Effective Regularization Methods for Training on Noisily Labeled Data with Generalization Guarantee
arXiv:1905.11368
Abstract
Over-parameterized deep neural networks trained by simple first-order methods are known to be able to fit any labeling of data. Such over-fitting ability hinders generalization when mislabeled training examples are present. On the other hand, simple regularization methods like early-stopping can often achieve highly nontrivial performance on clean test data in these scenarios, a phenomenon not theoretically understood. This paper proposes and analyzes two simple and intuitive regularization methods: (i) regularization by the distance between the network parameters to initialization, and (ii) adding a trainable auxiliary variable to the network output for each training example. Theoretically, we prove that gradient descent training with either of these two methods leads to a generalization guarantee on the clean data distribution despite being trained using noisy labels. Our generalization analysis relies on the connection between wide neural network and neural tangent kernel (NTK). The generalization bound is independent of the network size, and is comparable to the bound one can get when there is no label noise. Experimental results verify the effectiveness of these methods on noisily labeled datasets.
International Conference on Learning Representations (ICLR) 2020
References in corpus (19)
- Neural Tangent Kernel: Convergence and Generalization in Neural Networks
- Learning to Reweight Examples for Robust Deep Learning
- Classification with Noisy Labels by Importance Reweighting
- A Convergence Theory for Deep Learning via Over-Parameterization
- MentorNet: Learning Data-Driven Curriculum for Very Deep Neural Networks on Corrupted Labels
- Training Convolutional Networks with Noisy Labels
- Deep Learning is Robust to Massive Label Noise
- On Exact Computation with an Infinitely Wide Neural Net
- Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Early stopping and non-parametric regression: An optimal data-dependent stopping rule
- Scaling Limits of Wide Neural Networks with Weight Sharing: Gaussian Process Behavior, Gradient Independence, and Neural Tangent Kernel Derivation
- Learning and Generalization in Overparameterized Neural Networks, Going Beyond Two Layers
- Gradient Descent with Early Stopping is Provably Robust to Label Noise for Overparameterized Neural Networks
- Graph Neural Tangent Kernel: Fusing Graph Neural Networks with Graph Kernels
- Generalization in Deep Networks: The Role of Distance from Initialization
- Learning with Confident Examples: Rank Pruning for Robust Classification with Noisy Labels
- Generalization Error Bounds of Gradient Descent for Learning Over-parameterized Deep ReLU Networks
- Learning From Noisy Large-Scale Datasets With Minimal Supervision
Cited by in corpus (10)
- Open-set Label Noise Can Improve Robustness Against Inherent Label Noise
- Distillation Early Stopping? Harvesting Dark Knowledge Utilizing Anisotropic Information Retrieval For Overparameterized Neural Network
- Learning Not to Learn in the Presence of Noisy Labels
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural Networks
- Rethinking Influence Functions of Neural Networks in the Over-parameterized Regime
- Coresets for Robust Training of Neural Networks against Noisy Labels
- RATT: Leveraging Unlabeled Data to Guarantee Generalization
- An Exploration into why Output Regularization Mitigates Label Noise
- A Revision of Neural Tangent Kernel-based Approaches for Neural Networks
- Joint Text and Label Generation for Spoken Language Understanding