Variational Rectification Inference for Learning with Noisy Labels
arXiv:2603.17255 · doi:10.1007/s11263-024-02205-5
Abstract
Label noise has been broadly observed in real-world datasets. To mitigate the negative impact of overfitting to label noise for deep models, effective strategies (\textit{e.g.}, re-weighting, or loss rectification) have been broadly applied in prevailing approaches, which have been generally learned under the meta-learning scenario. Despite the robustness of noise achieved by the probabilistic meta-learning models, they usually suffer from model collapse that degenerates generalization performance. In this paper, we propose variational rectification inference (VRI) to formulate the adaptive rectification for loss functions as an amortized variational inference problem and derive the evidence lower bound under the meta-learning framework. Specifically, VRI is constructed as a hierarchical Bayes by treating the rectifying vector as a latent variable, which can rectify the loss of the noisy sample with the extra randomness regularization and is, therefore, more robust to label noise. To achieve the inference of the rectifying vector, we approximate its conditional posterior with an amortization meta-network. By introducing the variational term in VRI, the conditional posterior is estimated accurately and avoids collapsing to a Dirac delta function, which can significantly improve the generalization performance. The elaborated meta-network and prior network adhere to the smoothness assumption, enabling the generation of reliable rectification vectors. Given a set of clean meta-data, VRI can be efficiently meta-learned within the bi-level optimization programming. Besides, theoretical analysis guarantees that the meta-network can be efficiently learned with our algorithm. Comprehensive comparison experiments and analyses validate its effectiveness for robust learning with noisy labels, particularly in the presence of open-set noise.
References in corpus (27)
- Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence
- Generalized Cross Entropy Loss for Training Deep Neural Networks with Noisy Labels
- Classification with Noisy Labels by Importance Reweighting
- MixMatch: A Holistic Approach to Semi-Supervised Learning
- Training Deep Neural Networks on Noisy Labels with Bootstrapping
- MentorNet: Learning Data-Driven Curriculum for Very Deep Neural Networks on Corrupted Labels
- Training Convolutional Networks with Noisy Labels
- DivideMix: Learning with Noisy Labels as Semi-supervised Learning
- Meta-Weight-Net: Learning an Explicit Mapping For Sample Weighting
- Early-Learning Regularization Prevents Memorization of Noisy Labels
- Unsupervised Label Noise Modeling and Loss Correction
- Toward Robustness against Label Noise in Training Deep Discriminative Neural Networks
- Using Trusted Data to Train Deep Networks on Labels Corrupted by Severe Noise
- Bilevel Programming for Hyperparameter Optimization and Meta-Learning
- Dual T: Reducing Estimation Error for Transition Matrix in Label-noise Learning
- Peer Loss Functions: Learning from Noisy Labels without Knowing Noise Rates
- Understanding and Improving Early Stopping for Learning with Noisy Labels
- MoPro: Webly Supervised Learning with Momentum Prototypes
- Instance-dependent Label-noise Learning under a Structural Causal Model
- Robust Training under Label Noise by Over-parameterization
- Learning with Instance-Dependent Label Noise: A Sample Sieve Approach
- Learning Noise Transition Matrix from Only Noisy Labels via Total Variation Regularization
- Generalized Jensen-Shannon Divergence Loss for Learning with Noisy Labels
- Meta-Learning with Shared Amortized Variational Inference
- Robust Probabilistic Modeling with Bayesian Data Reweighting
- Stability and Generalization of Bilevel Programming in Hyperparameter Optimization