Learning with Bounded Instance- and Label-dependent Label Noise
arXiv:1709.03768
Abstract
Instance- and Label-dependent label Noise (ILN) widely exists in real-world datasets but has been rarely studied. In this paper, we focus on Bounded Instance- and Label-dependent label Noise (BILN), a particular case of ILN where the label noise rates -- the probabilities that the true labels of examples flip into the corrupted ones -- have upper bound less than . Specifically, we introduce the concept of distilled examples, i.e. examples whose labels are identical with the labels assigned for them by the Bayes optimal classifier, and prove that under certain conditions classifiers learnt on distilled examples will converge to the Bayes optimal classifier. Inspired by the idea of learning with distilled examples, we then propose a learning algorithm with theoretical guarantees for its robustness to BILN. At last, empirical evaluations on both synthetic and real-world datasets show effectiveness of our algorithm in learning with BILN.
Published in the International Conference on Machine Learning (ICML), 2020
References in corpus (10)
- Learning to Reweight Examples for Robust Deep Learning
- Dimensionality-Driven Learning with Noisy Labels
- Masking: A New Perspective of Noisy Supervision
- Joint Optimization Framework for Learning with Noisy Labels
- Learning from Complementary Labels
- Learning from Noisy Labels with Distillation
- Learning with Confident Examples: Rank Pruning for Robust Classification with Noisy Labels
- Iterative Learning with Open-set Noisy Labels
- Revisiting Perceptron: Efficient and Label-Optimal Learning of Halfspaces
- Efficient Learning of Linear Separators under Bounded Noise
Cited by in corpus (6)
- Instance-Dependent PU Learning by Bayesian Optimal Relabeling
- Heteroskedastic and Imbalanced Deep Learning with Adaptive Regularization
- Learning with Biased Complementary Labels
- Multi-Modal Multi-Scale Deep Learning for Large-Scale Image Annotation
- Multi-Class Classification from Noisy-Similarity-Labeled Data
- A Generalized Neyman-Pearson Criterion for Optimal Domain Adaptation