Learning with Average Top-k Loss
arXiv:1705.08826
Abstract
In this work, we introduce the {\em average top-} (\atk) loss as a new aggregate loss for supervised learning, which is the average over the largest individual losses over a training dataset. We show that the \atk loss is a natural generalization of the two widely used aggregate losses, namely the average loss and the maximum loss, but can combine their advantages and mitigate their drawbacks to better adapt to different data distributions. Furthermore, it remains a convex function over all individual losses, which can lead to convex optimization problems that can be solved effectively with conventional gradient-based methods. We provide an intuitive interpretation of the \atk loss based on its equivalent effect on the continuous individual loss functions, suggesting that it can reduce the penalty on correctly classified data. We further give a learning theory analysis of \matk learning on the classification calibration of the \atk loss and the error bounds of \atk-SVM. We demonstrate the applicability of minimum average top- learning for binary classification and regression using synthetic and real datasets.
18 pages
Cited by in corpus (22)
- Long-tail learning via logit adjustment
- Large-Scale Methods for Distributionally Robust Optimization
- Fairness risk measures
- Optimal Epoch Stochastic Gradient Descent Ascent Methods for Min-Max Optimization
- Learning Bounds for Risk-sensitive Learning
- Geometry-Inspired Top-k Adversarial Perturbations
- Adaptive Sampling for Stochastic Risk-Averse Learning
- Learning by Minimizing the Sum of Ranked Range
- On Tilted Losses in Machine Learning: Theory and Applications
- When Do Curricula Work?
- Coping with Label Shift via Distributionally Robust Optimisation
- Spectral risk-based learning using unbounded losses
- A Simple and Effective Framework for Pairwise Deep Metric Learning
- Are Negative Samples Necessary in Entity Alignment? An Approach with High Performance, Scalability and Robustness
- TML-AP: Adversarial Attacks to Top- Multi-Label Learning
- Doubly-stochastic mining for heterogeneous retrieval
- PrimA6D: Rotational Primitive Reconstruction for Enhanced and Robust 6D Pose Estimation
- Efficient Online-Bandit Strategies for Minimax Learning Problems
- Learning Neural Models for Natural Language Processing in the Face of Distributional Shift
- A Univariate Bound of Area Under ROC
- Deep Metric Learning with Locality Sensitive Angular Loss for Self-Correcting Source Separation of Neural Spiking Signals
- Sum of Ranked Range Loss for Supervised Learning