machine learning

PUe: Biased Positive-Unlabeled Learning Enhancement by Causal Inference

arXiv:2607.13428

summary

The paper proposes PUe, a framework that improves positive‑unlabeled (PU) learning under biased label selection by using normalized propensity scores and inverse probability weighting, providing new risk formulations and theoretical analysis, and demonstrating gains on image and medical datasets.

Abstract

Positive-Unlabeled (PU) learning aims to achieve high-accuracy binary classification with limited labeled positive examples and numerous unlabeled ones. Existing cost-sensitive-based methods often rely on strong assumptions that examples with an observed positive label were selected entirely at random. In fact, the uneven distribution of labels is prevalent in real-world PU problems, indicating that most actual positive and unlabeled data are subject to selection bias. Building on the SAR-PU propensity-weighted framework of Bekker et al., we study a PU learning enhancement (PUe) framework using normalized propensity scores and normalized inverse probability weighting (NIPW). PUe's main contributions are a normalized inverse-probability-weighted PU risk formulation; additional theoretical analyses of normalized sample-weight error and common PU estimators under biased labeling; regularized deep propensity-score estimation; integration with modern cost-sensitive PU methods; and support for selectively labeled negative classes. Experiments on MNIST, CIFAR-10, and ADNI demonstrate improvements over several PU baselines under non-uniform label distributions.

Extended arXiv version of the NeurIPS 2023 paper; includes additional discussion of related SAR-PU work

Topics & keywords

#positive-unlabeled learning#selection bias#propensity scoring#causal inference#deep learning#risk estimationPU learningpropensity scoreinverse probability weightingbiased labelingnormalized riskSAR-PU