On surrogate loss functions and -divergences
arXiv:math/0510521 · doi:10.1214/08-AOS595
Abstract
The goal of binary classification is to estimate a discriminant function from observations of covariate vectors and corresponding binary labels. We consider an elaboration of this problem in which the covariates are not available directly but are transformed by a dimensionality-reducing quantizer . We present conditions on loss functions such that empirical risk minimization yields Bayes consistency when both the discriminant function and the quantizer are estimated. These conditions are stated in terms of a general correspondence between loss functions and a class of functionals known as Ali-Silvey or -divergence functionals. Whereas this correspondence was established by Blackwell [Proc. 2nd Berkeley Symp. Probab. Statist. 1 (1951) 93--102. Univ. California Press, Berkeley] for the 0--1 loss, we extend the correspondence to the broader class of surrogate loss functions that play a key role in the general theory of Bayes consistency for binary classification. Our result makes it possible to pick out the (strict) subset of surrogate loss functions that yield Bayes consistency for joint estimation of the discriminant function and the quantizer.
Published in at http://dx.doi.org/10.1214/08-AOS595 the Annals of Statistics (http://www.imstat.org/aos/) by the Institute of Mathematical Statistics (http://www.imstat.org)
Cited by in corpus (10)
- Estimating divergence functionals and the likelihood ratio by convex risk minimization
- The potential and perils of preprocessing: Building new foundations
- On distributionally robust extreme value analysis
- Minimax Robust Detection: Classic Results and Recent Advances
- Direct estimation of density functionals using a polynomial basis
- Understanding Compressive Adversarial Privacy
- Proximity Operators of Discrete Information Divergences
- Information measures and geometry of the hyperbolic exponential families of Poincaré and hyperboloid distributions
- On the Minimization of Convex Functionals of Probability Distributions Under Band Constraints
- An Algorithm for Learning Smaller Representations of Models With Scarce Data