Pointwise Binary Classification with Pairwise Confidence Comparisons
arXiv:2010.01875
Abstract
To alleviate the data requirement for training effective binary classifiers in binary classification, many weakly supervised learning settings have been proposed. Among them, some consider using pairwise but not pointwise labels, when pointwise labels are not accessible due to privacy, confidentiality, or security reasons. However, as a pairwise label denotes whether or not two data points share a pointwise label, it cannot be easily collected if either point is equally likely to be positive or negative. Thus, in this paper, we propose a novel setting called pairwise comparison (Pcomp) classification, where we have only pairs of unlabeled data that we know one is more likely to be positive than the other. Firstly, we give a Pcomp data generation process, derive an unbiased risk estimator (URE) with theoretical guarantee, and further improve URE using correction functions. Secondly, we link Pcomp classification to noisy-label learning to develop a progressive URE and improve it by imposing consistency regularization. Finally, we demonstrate by experiments the effectiveness of our methods, which suggests Pcomp is a valuable and practically useful type of pairwise supervision besides the pairwise label.
Accepted to ICML 2021
References in corpus (12)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results
- Are Anchor Points Really Indispensable in Label-Noise Learning?
- Learning from Complementary Labels
- Learning with Confident Examples: Rank Pruning for Robust Classification with Noisy Labels
- Semi-Supervised AUC Optimization based on Positive-Unlabeled Learning
- On the Minimal Supervision for Training Any Binary Classifier from Only Unlabeled Data
- Mitigating Overfitting in Supervised Classification from Two Unlabeled Datasets: A Consistent Risk Correction Approach
- Self-PU: Self Boosted and Calibrated Positive-Unlabeled Training
- Uncoupled Regression from Pairwise Comparison Data
- Regression with Comparisons: Escaping the Curse of Dimensionality with Ordinal Information