Does Distributionally Robust Supervised Learning Give Robust Classifiers?
arXiv:1611.02041
Abstract
Distributionally Robust Supervised Learning (DRSL) is necessary for building reliable machine learning systems. When machine learning is deployed in the real world, its performance can be significantly degraded because test data may follow a different distribution from training data. DRSL with f-divergences explicitly considers the worst-case distribution shift by minimizing the adversarially reweighted training loss. In this paper, we analyze this DRSL, focusing on the classification scenario. Since the DRSL is explicitly formulated for a distribution shift scenario, we naturally expect it to give a robust classifier that can aggressively handle shifted distributions. However, surprisingly, we prove that the DRSL just ends up giving a classifier that exactly fits the given training distribution, which is too pessimistic. This pessimism comes from two sources: the particular losses used in classification and the fact that the variety of distributions to which the DRSL tries to be robust is too wide. Motivated by our analysis, we propose simple DRSL that overcomes this pessimism and empirically demonstrate its effectiveness.
ICML 2018 camera-ready (final submission version)
Cited by in corpus (56)
- A review of domain adaptation without target labels
- Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization
- WILDS: A Benchmark of in-the-Wild Distribution Shifts
- Learning Robust Global Representations by Penalizing Local Predictive Power
- Out-of-Distribution Generalization via Risk Extrapolation (REx)
- Towards Out-Of-Distribution Generalization: A Survey
- Distributionally Robust Optimization: A Review
- Curriculum Loss: Robust Learning and Generalization against Label Corruption
- An Investigation of Why Overparameterization Exacerbates Spurious Correlations
- Out of Distribution Generalization in Machine Learning
- Sample Out-Of-Sample Inference Based on Wasserstein Distance
- BREEDS: Benchmarks for Subpopulation Shift
- Unlearn Dataset Bias in Natural Language Inference by Fitting the Residual
- Gradient Matching for Domain Generalization
- Classification from Positive, Unlabeled and Biased Negative Data
- Label-Imbalanced and Group-Sensitive Classification under Overparameterization
- Calibrated Surrogate Losses for Adversarially Robust Classification
- Algorithmic Bias and Data Bias: Understanding the Relation between Distributionally Robust Optimization and Data Curation
- Kernelized Heterogeneous Risk Minimization
- Distributionally Robust Deep Learning using Hardness Weighted Sampling
- Representation Matters: Assessing the Importance of Subgroup Allocations in Training Data
- Support Vector Machine Classifier via Soft-Margin Loss
- Coping with Label Shift via Distributionally Robust Optimisation
- DORO: Distributional and Outlier Robust Optimization
- Model-Based Domain Generalization
- An Online Learning Approach to Interpolation and Extrapolation in Domain Generalization
- Evaluating Robustness to Dataset Shift via Parametric Robustness Sets
- Estimating the Brittleness of AI: Safety Integrity Levels and the Need for Testing Out-Of-Distribution Performance
- Toward Learning Human-aligned Cross-domain Robust Models by Countering Misaligned Features
- Learning Representations that Support Robust Transfer of Predictors
- Non-convex Distributionally Robust Optimization: Non-asymptotic Analysis
- Zero-shot Domain Adaptation Based on Attribute Information
- Pulling Up by the Causal Bootstraps: Causal Data Augmentation for Pre-training Debiasing
- Less Is Better: Unweighted Data Subsampling via Influence Function
- Focus on the Common Good: Group Distributional Robustness Follows
- Modeling the Second Player in Distributionally Robust Optimization
- Algorithm Fairness in AI for Medicine and Healthcare
- Predict then Interpolate: A Simple Algorithm to Learn Stable Classifiers
- Coordinate Linear Variance Reduction for Generalized Linear Programming
- Towards Robust Off-Policy Evaluation via Human Inputs
- Learning Stable Classifiers by Transferring Unstable Features
- Robust binary classification with the 01 loss
- Unsupervised Learning of Debiased Representations with Pseudo-Attributes
- A call for better unit testing for invariant risk minimisation
- Learning Neural Models for Natural Language Processing in the Face of Distributional Shift
- Adversarial Regression with Doubly Non-negative Weighting Matrices
- Learning Under Adversarial and Interventional Shifts
- Explainability-aided Domain Generalization for Image Classification
- Principled learning method for Wasserstein distributionally robust optimization with local perturbations
- Towards adversarial robustness with 01 loss neural networks
- Boosted CVaR Classification
- Robust Learning in Heterogeneous Contexts
- Model Specification Test with Unlabeled Data: Approach from Covariate Shift
- Distributionally Robust Learning with Stable Adversarial Training
- Zeroth-Order Methods for Convex-Concave Minmax Problems: Applications to Decision-Dependent Risk Minimization
- Distributionally Robust Language Modeling