Nonlinear Invariant Risk Minimization: A Causal Approach
arXiv:2102.12353
Abstract
Due to spurious correlations, machine learning systems often fail to generalize to environments whose distributions differ from the ones used at training time. Prior work addressing this, either explicitly or implicitly, attempted to find a data representation that has an invariant relationship with the target. This is done by leveraging a diverse set of training environments to reduce the effect of spurious features and build an invariant predictor. However, these methods have generalization guarantees only when both data representation and classifiers come from a linear model class. We propose invariant Causal Representation Learning (iCaRL), an approach that enables out-of-distribution (OOD) generalization in the nonlinear setting (i.e., nonlinear representations and nonlinear classifiers). It builds upon a practical and general assumption: the prior over the data representation (i.e., a set of latent variables encoding the data) given the target and the environment belongs to general exponential family distributions. Based on this, we show that it is possible to identify the data representation up to simple transformations. We also prove that all direct causes of the target can be fully discovered, which further enables us to obtain generalization guarantees in the nonlinear setting. Extensive experiments on both synthetic and real-world datasets show that our approach outperforms a variety of baseline methods. Finally, in the discussion, we further explore the aforementioned assumption and propose a more general hypothesis, called the Agnostic Hypothesis: there exist a set of hidden causal factors affecting both inputs and outcomes. The Agnostic Hypothesis can provide a unifying view of machine learning. More importantly, it can inspire a new direction to explore a general theory for identifying hidden causal factors, which is key to enabling the OOD generalization guarantees.
References in corpus (17)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Domain Generalization via Invariant Feature Representation
- Domain Generalization via Model-Agnostic Learning of Semantic Features
- Kernel-based Conditional Independence Test and Application in Causal Discovery
- Improve Unsupervised Domain Adaptation with Mixup Training
- Inferring deterministic causal relations
- Learning Robust Representations by Projecting Superficial Statistics Out
- Identifying the consequences of dynamic treatment strategies: A decision-theoretic overview
- The Risks of Invariant Risk Minimization
- In Search of Lost Domain Generalization
- Feature-Critic Networks for Heterogeneous Domain Generalization
- Understanding the Failure Modes of Out-of-Distribution Generalization
- Out of Distribution Generalization in Machine Learning
- Representation Learning via Invariant Causal Mechanisms
- Deconfounding Reinforcement Learning in Observational Settings
- Empirical or Invariant Risk Minimization? A Sample Complexity Perspective
- Latent Causal Invariant Model
Cited by in corpus (9)
- Self-Supervised Learning with Data Augmentations Provably Isolates Content from Style
- Invariance Principle Meets Information Bottleneck for Out-of-Distribution Generalization
- Model-Based Domain Generalization
- Visual Representation Learning Does Not Generalize Strongly Within the Same Domain
- Environment Invariant Linear Least Squares
- Towards Principled Disentanglement for Domain Generalization
- Optimization-based Causal Estimation from Heterogenous Environments
- Contrastive ACE: Domain Generalization Through Alignment of Causal Mechanisms
- Invariant Risk Minimisation for Cross-Organism Inference: Substituting Mouse Data for Human Data in Human Risk Factor Discovery