Loss factorization, weakly supervised learning and label noise robustness
arXiv:1602.02450
Abstract
We prove that the empirical risk of most well-known loss functions factors into a linear term aggregating all labels with a term that is label free, and can further be expressed by sums of the loss. This holds true even for non-smooth, non-convex losses and in any RKHS. The first term is a (kernel) mean operator --the focal quantity of this work-- which we characterize as the sufficient statistic for the labels. The result tightens known generalization bounds and sheds new light on their interpretation. Factorization has a direct application on weakly supervised learning. In particular, we demonstrate that algorithms like SGD and proximal methods can be adapted with minimal effort to handle weak supervision, once the mean operator has been estimated. We apply this idea to learning with asymmetric noisy labels, connecting and extending prior work. Furthermore, we show that most losses enjoy a data-dependent (by the mean operator) form of noise robustness, in contrast with known negative results.
References in corpus (2)
Cited by in corpus (16)
- Private federated learning on vertically partitioned data via entity resolution and additively homomorphic encryption
- Image Classification with Deep Learning in the Presence of Noisy Labels: A Survey
- Positive-Unlabeled Learning with Non-Negative Risk Estimator
- Decoupling "when to update" from "how to update"
- Making Deep Neural Networks Robust to Label Noise: a Loss Correction Approach
- Learning Graph Neural Networks with Noisy Labels
- Learning from Binary Labels with Instance-Dependent Corruption
- A Semi-Supervised Two-Stage Approach to Learning from Noisy Labels
- Calibrated Surrogate Losses for Adversarially Robust Classification
- Principled analytic classifier for positive-unlabeled learning via weighted integral probability metric
- Machine learning approach for the search of resonances with topological features at the Large Hadron Collider
- Label Propagation for Learning with Label Proportions
- Do We Really Need Gold Samples for Sample Weighting Under Label Noise?
- Exploiting Class Similarity for Machine Learning with Confidence Labels and Projective Loss Functions
- Learning with Inadequate and Incorrect Supervision
- The Crossover Process: Learnability and Data Protection from Inference Attacks