Regularization via Mass Transportation
arXiv:1710.10016
Abstract
The goal of regression and classification methods in supervised learning is to minimize the empirical risk, that is, the expectation of some loss function quantifying the prediction error under the empirical distribution. When facing scarce training data, overfitting is typically mitigated by adding regularization terms to the objective that penalize hypothesis complexity. In this paper we introduce new regularization techniques using ideas from distributionally robust optimization, and we give new probabilistic interpretations to existing techniques. Specifically, we propose to minimize the worst-case expected loss, where the worst case is taken over the ball of all (continuous or discrete) distributions that have a bounded transportation distance from the (discrete) empirical distribution. By choosing the radius of this ball judiciously, we can guarantee that the worst-case expected loss provides an upper confidence bound on the loss on test data, thus offering new generalization bounds. We prove that the resulting regularized learning problems are tractable and can be tractably kernelized for many popular loss functions. We validate our theoretical out-of-sample guarantees through simulated and empirical experiments.
Cited by in corpus (39)
- Robust Wasserstein Profile Inference and Applications to Machine Learning
- Towards Out-Of-Distribution Generalization: A Survey
- Distributionally Robust Optimization: A Review
- Robust Federated Learning: The Case of Affine Distribution Shifts
- Robust Hypothesis Testing Using Wasserstein Uncertainty Sets
- Wasserstein Distributionally Robust Kalman Filtering
- Optimal Transport Based Distributionally Robust Optimization: Structural Properties and Iterative Schemes
- Sensitivity analysis of Wasserstein distributionally robust optimization problems
- A Distributionally Robust Approach to Fair Classification
- Prescribing net demand for two-stage electricity generation scheduling
- A Distributionally Robust Area Under Curve Maximization Model
- Distributionally Robust Local Non-parametric Conditional Estimation
- A First-Order Algorithmic Framework for Wasserstein Distributionally Robust Logistic Regression
- Partition-based Distributionally Robust Optimization via Optimal Transport with Order Cone Constraints
- Sinkhorn Distributionally Robust Optimization
- Fast Distributionally Robust Learning with Variance Reduced Min-Max Optimization
- Imbalanced Image Classification with Complement Cross Entropy
- Fast Epigraphical Projection-based Incremental Algorithms for Wasserstein Distributionally Robust Support Vector Machine
- Distributionally Robust Graph Learning from Smooth Signals under Moment Uncertainty
- Lipschitz Networks and Distributional Robustness
- Generalised Lipschitz Regularisation Equals Distributional Robustness
- Minimax Classification with 0-1 Loss and Performance Guarantees
- Online Optimization and Ambiguity-based Learning of Distributionally Uncertain Dynamic Systems
- Distributionally Robust Formulation and Model Selection for the Graphical Lasso
- Distributionally Robust Parametric Maximum Likelihood Estimation
- General Supervision via Probabilistic Transformations
- On the Impossibility of Statistically Improving Empirical Optimization: A Second-Order Stochastic Dominance Perspective
- Distributional Robustness with IPMs and links to Regularization and GANs
- Distributionally Robust Weighted -Nearest Neighbors
- Adversarially Robust Kernel Smoothing
- Robust Generalization despite Distribution Shift via Minimum Discriminating Information
- Distributionally Robust Multi-Output Regression Ranking
- Statistical Robustness of Empirical Risks in Machine Learning
- A Robust Learning Algorithm for Regression Models Using Distributionally Robust Optimization under the Wasserstein Metric
- Principled learning method for Wasserstein distributionally robust optimization with local perturbations
- Group-Structured Adversarial Training
- Orthounimodal Distributionally Robust Optimization: Representation, Computation and Multivariate Extreme Event Applications
- Higher-Order Expansion and Bartlett Correctability of Distributionally Robust Optimization
- Efficient Data-Driven Optimization with Noisy Data