Causally motivated Shortcut Removal Using Auxiliary Labels
arXiv:2105.06422
Abstract
Shortcut learning, in which models make use of easy-to-represent but unstable associations, is a major failure mode for robust machine learning. We study a flexible, causally-motivated approach to training robust predictors by discouraging the use of specific shortcuts, focusing on a common setting where a robust predictor could achieve optimal \emph{iid} generalization in principle, but is overshadowed by a shortcut predictor in practice. Our approach uses auxiliary labels, typically available at training time, to enforce conditional independences implied by the causal graph. We show both theoretically and empirically that causally-motivated regularization schemes (a) lead to more robust estimators that generalize well under distribution shift, and (b) have better finite sample efficiency compared to usual regularization schemes, even when no shortcut is present. Our analysis highlights important theoretical properties of training techniques commonly used in the causal inference, fairness, and disentanglement literatures. Our code is available at https://github.com/mymakar/causally_motivated_shortcut_removal
References in corpus (10)
- Learning Transferable Features with Deep Adaptation Networks
- Deep Domain Confusion: Maximizing for Domain Invariance
- Shortcut Learning in Deep Neural Networks
- Underspecification Presents Challenges for Credibility in Modern Machine Learning
- Improving Robustness Without Sacrificing Accuracy with Patch Gaussian Augmentation
- The Risks of Invariant Risk Minimization
- An Investigation of Why Overparameterization Exacerbates Spurious Correlations
- Understanding the Failure Modes of Out-of-Distribution Generalization
- On the Benefits of Invariance in Neural Networks
- Out-of-distribution Prediction with Invariant Risk Minimization: The Limitation and An Effective Fix