A Survey of Label-noise Representation Learning: Past, Present and Future
arXiv:2011.04406
Abstract
Classical machine learning implicitly assumes that labels of the training data are sampled from a clean distribution, which can be too restrictive for real-world scenarios. However, statistical-learning-based methods may not train deep learning models robustly with these noisy labels. Therefore, it is urgent to design Label-Noise Representation Learning (LNRL) methods for robustly training deep models with noisy labels. To fully understand LNRL, we conduct a survey study. We first clarify a formal definition for LNRL from the perspective of machine learning. Then, via the lens of learning theory and empirical study, we figure out why noisy labels affect deep models' performance. Based on the theoretical guidance, we categorize different LNRL methods into three directions. Under this unified taxonomy, we provide a thorough discussion of the pros and cons of different categories. More importantly, we summarize the essential components of robust LNRL, which can spark new directions. Lastly, we propose possible research directions within LNRL, such as new datasets, instance-dependent LNRL, and adversarial LNRL. We also envision potential directions beyond LNRL, such as learning with feature-noise, preference-noise, domain-noise, similarity-noise, graph-noise and demonstration-noise.
The draft is kept updating; any comments and suggestions are welcome
References in corpus (9)
- DivideMix: Learning with Noisy Labels as Semi-supervised Learning
- Unsupervised Label Noise Modeling and Loss Correction
- How does Disagreement Help Generalization against Label Corruption?
- Understanding and Utilizing Deep Neural Networks Trained with Noisy Labels
- A more robust boosting algorithm
- Maximum Entropy Semi-Supervised Inverse Reinforcement Learning
- Towards Robust ResNet: A Small Step but A Giant Leap
- A Second-Order Approach to Learning with Instance-Dependent Label Noise
- Multi-Class Classification from Noisy-Similarity-Labeled Data
Cited by in corpus (5)
- Unsupervised Domain Adaptation of Black-Box Source Models
- Extended T: Learning with Mixed Closed-set and Open-set Noisy Labels
- Learning From Long-Tailed Data With Noisy Labels
- A Second-Order Approach to Learning with Instance-Dependent Label Noise
- Knowledge Distillation with Noisy Labels for Natural Language Understanding