Learning with Confident Examples: Rank Pruning for Robust Classification with Noisy Labels
arXiv:1705.01936
Abstract
Noisy PN learning is the problem of binary classification when training examples may be mislabeled (flipped) uniformly with noise rate rho1 for positive examples and rho0 for negative examples. We propose Rank Pruning (RP) to solve noisy PN learning and the open problem of estimating the noise rates, i.e. the fraction of wrong positive and negative labels. Unlike prior solutions, RP is time-efficient and general, requiring O(T) for any unrestricted choice of probabilistic classifier with T fitting time. We prove RP has consistent noise estimation and equivalent expected risk as learning with uncorrupted labels in ideal conditions, and derive closed-form solutions when conditions are non-ideal. RP achieves state-of-the-art noise estimation and F1, error, and AUC-PR for both MNIST and CIFAR datasets, regardless of the amount of noise and performs similarly impressively when a large portion of training examples are noise drawn from a third distribution. To highlight, RP with a CNN classifier can predict if an MNIST digit is a "one"or "not" with only 0.25% error, and 0.46 error across all digits, even when 50% of positive examples are mislabeled and 50% of observed positive labels are mislabeled negative examples.
References in corpus (3)
Cited by in corpus (23)
- Generalized Cross Entropy Loss for Training Deep Neural Networks with Noisy Labels
- Learning from positive and unlabeled data: a survey
- Are Anchor Points Really Indispensable in Label-Noise Learning?
- Learning to detect chest radiographs containing lung nodules using visual attention networks
- Active Bias: Training More Accurate Neural Networks by Emphasizing High Variance Samples
- Dual T: Reducing Estimation Error for Transition Matrix in Label-noise Learning
- Learning with Bounded Instance- and Label-dependent Label Noise
- Provably End-to-end Label-Noise Learning without Anchor Points
- Generalized Jensen-Shannon Divergence Loss for Learning with Noisy Labels
- Clusterability as an Alternative to Anchor Points When Learning with Noisy Labels
- Simple and Effective Regularization Methods for Training on Noisily Labeled Data with Generalization Guarantee
- Fair Classification with Noisy Protected Attributes: A Framework with Provable Guarantees
- Robust and On-the-fly Dataset Denoising for Image Classification
- Intelligence, physics and information -- the tradeoff between accuracy and simplicity in machine learning
- NLNL: Negative Learning for Noisy Labels
- Pointwise Binary Classification with Pairwise Confidence Comparisons
- Robust Deep Learning with Active Noise Cancellation for Spatial Computing
- Joint Negative and Positive Learning for Noisy Labels
- Consistency Regularization Can Improve Robustness to Label Noise
- A Non-Intrusive Correction Algorithm for Classification Problems with Corrupted Data
- Mitigating Class Boundary Label Uncertainty to Reduce Both Model Bias and Variance
- Semi-Supervised Domain Adaptation via Selective Pseudo Labeling and Progressive Self-Training
- Generating Relevant Counter-Examples from a Positive Unlabeled Dataset for Image Classification