machine learning

Generalized Distribution-Free Semi-Supervised Learning with Risk Rewrite

arXiv:2607.11947

summary

The paper introduces a generalized, distribution‑free framework for semi‑supervised learning that builds unbiased risk estimators for both binary and multiclass problems, achieving lower variance than previous methods and providing theoretical and empirical performance gains.

Abstract

Typical semi-supervised learning (SSL) methods rely on distributional assumptions, and their performance degrades when these are violated. While PNU learning, a risk rewriting method, offers a distribution-free alternative, it is restricted to binary classification and its variance optimality remains unclear. In this paper, we propose a generalized framework that constructs unbiased risk estimators using linear combinations of component risks, subsuming PNU learning and extending to multiclass classification. We derive the minimum achievable variance, demonstrating our estimator can attain lower variance than PNU in asymmetric loss scenarios. Furthermore, we establish a generalization bound directly linking this variance reduction to improved learning performance. Based on these theoretical insights, we introduce two practical SSL methods that empirically match or outperform existing approaches on binary and multiclass benchmarks.

Accepted to The Conference on Uncertainty in Artificial Intelligence (UAI) 2026

Topics & keywords

#semi-supervised learning#distribution-free learning#multiclass classification#risk estimation#variance reductionrisk rewritePNU learningunbiased risk estimatorgeneralization boundasymmetric loss