A Survey of Learning Criteria Going Beyond the Usual Risk
arXiv:2110.04996 · doi:10.1613/jair.1.15000
Abstract
Virtually all machine learning tasks are characterized using some form of loss function, and "good performance" is typically stated in terms of a sufficiently small average loss, taken over the random draw of test data. While optimizing for performance on average is intuitive, convenient to analyze in theory, and easy to implement in practice, such a choice brings about trade-offs. In this work, we survey and introduce a wide variety of non-traditional criteria used to design and evaluate machine learning algorithms, place the classical paradigm within the proper historical context, and propose a view of learning problems which emphasizes the question of "what makes for a desirable loss distribution?" in place of tacit use of the expected loss.
Final version published in JAIR
References in corpus (20)
- Empirical Bernstein Bounds and Sample Variance Penalization
- Risk-Sensitive and Robust Decision-Making: a CVaR Optimization Approach
- Deep Reinforcement Learning at the Edge of the Statistical Precipice
- Fairness risk measures
- Statistical Learning with Conditional Value at Risk
- Learning Bounds for Risk-sensitive Learning
- Adaptive Sampling for Stochastic Risk-Averse Learning
- On Tilted Losses in Machine Learning: Theory and Applications
- Probabilistically Robust Learning: Balancing Average- and Worst-case Performance
- DORO: Distributional and Outlier Robust Optimization
- Spectral risk-based learning using unbounded losses
- Learning with risks based on M-location
- Risk-Adaptive Approaches to Stochastic Optimization: A Survey
- Supervised Learning with General Risk Functionals
- Tailoring to the Tails: Risk Measures for Fine-Grained Tail Sensitivity
- Unifying Lower Bounds on Prediction Dimension of Consistent Convex Surrogates
- Rank-based Decomposable Losses in Machine Learning: A Survey
- Metric Elicitation; Moving from Theory to Practice
- Flexible risk design using bi-directional dispersion
- Robust Generalization despite Distribution Shift via Minimum Discriminating Information