Composite Binary Losses
arXiv:0912.3301
Abstract
We study losses for binary classification and class probability estimation and extend the understanding of them from margin losses to general composite losses which are the composition of a proper loss with a link function. We characterise when margin losses can be proper composite losses, explicitly show how to determine a symmetric loss in full from half of one of its partial losses, introduce an intrinsic parametrisation of composite binary losses and give a complete characterisation of the relationship between proper losses and ``classification calibrated'' losses. We also consider the question of the ``best'' surrogate binary loss. We introduce a precise notion of ``best'' and show there exist situations where two convex surrogate losses are incommensurable. We provide a complete explicit characterisation of the convexity of composite binary losses in terms of the link function and the weight function associated with the proper loss which make up the composite loss. This characterisation suggests new ways of ``surrogate tuning''. Finally, in an appendix we present some new algorithm-independent results on the relationship between properness, convexity and robustness to misclassification noise for binary losses and show that all convex proper losses are non-robust to misclassification noise.
38 pages, 4 figures. Submitted to JMLR
References in corpus (1)
Cited by in corpus (51)
- Making Deep Neural Networks Robust to Label Noise: a Loss Correction Approach
- Entity Resolution and Federated Learning get a Federated Resolution
- Long-tail learning via logit adjustment
- Generative Adversarial Nets from a Density Ratio Estimation Perspective
- On the Minimal Supervision for Training Any Binary Classifier from Only Unlabeled Data
- Surrogate Regret Bounds for Bipartite Ranking via Strongly Proper Losses
- Loss factorization, weakly supervised learning and label noise robustness
- Confidence Scores Make Instance-dependent Label-noise Learning Possible
- On the Consistency of Ordinal Regression Methods
- Multiclass Learning with Simplex Coding
- Transfer Learning with Label Noise
- Imitation Learning from Imperfect Demonstration
- Convex Calibration Dimension for Multiclass Loss Matrices
- Multiclass Classification Calibration Functions
- Calibrated Surrogate Losses for Adversarially Robust Classification
- Revisiting One-vs-All Classifiers for Predictive Uncertainty and Out-of-Distribution Detection in Neural Networks
- Learning in the Presence of Corruption
- Learning to Abstain from Binary Prediction
- A General Theory for Structured Prediction with Smooth Convex Surrogates
- Label-Noise Robust Generative Adversarial Networks
- Exp-Concavity of Proper Composite Losses
- Learning from Indirect Observations
- A scaled Bregman theorem with applications
- Geometric Losses for Distributional Learning
- Bregman Divergence Bounds and Universality Properties of the Logarithmic Loss
- Permutation Weighting
- All your loss are belong to Bayes
- Threshold Choice Methods: the Missing Link
- Kernel Robust Bias-Aware Prediction under Covariate Shift
- Supervised Learning: No Loss No Cry
- Boosted and Differentially Private Ensembles of Decision Trees
- On Focal Loss for Class-Posterior Probability Estimation: A Theoretical Perspective
- General Probabilistic Surface Optimization and Log Density Estimation
- Cost Sensitive Learning in the Presence of Symmetric Label Noise
- Regret Bounds for Non-decomposable Metrics with Missing Labels
- Surrogate regret bounds for generalized classification performance metrics
- Combining Forecasts Using Ensemble Learning
- Scalable Approximations for Generalized Linear Problems
- Sparse Continuous Distributions and Fenchel-Young Losses
- The Convexity and Design of Composite Multiclass Losses
- Positive-Unlabeled Classification under Class Prior Shift and Asymmetric Error
- Dual Stochastic Natural Gradient Descent and convergence of interior half-space gradient approximations
- Robustness and Reliability When Training With Noisy Labels
- Strictly Proper Kernel Scoring Rules and Divergences with an Application to Kernel Two-Sample Hypothesis Testing
- Refinement revisited with connections to Bayes error, conditional entropy and calibrated classifiers
- Bregman-divergence-guided Legendre exponential dispersion model with finite cumulants (K-LED)
- Sum of Ranked Range Loss for Supervised Learning
- Evolving a Vector Space with any Generating Set
- A Symmetric Loss Perspective of Reliable Machine Learning
- A Hybrid Loss for Multiclass and Structured Prediction
- Contrastive Pessimistic Likelihood Estimation for Semi-Supervised Classification