Bayesian Nonparametric Causal Inference: Information Rates and Learning Algorithms
arXiv:1712.08914 · doi:10.1109/JSTSP.2018.2848230
Abstract
We investigate the problem of estimating the causal effect of a treatment on individual subjects from observational data, this is a central problem in various application domains, including healthcare, social sciences, and online advertising. Within the Neyman Rubin potential outcomes model, we use the Kullback Leibler (KL) divergence between the estimated and true distributions as a measure of accuracy of the estimate, and we define the information rate of the Bayesian causal inference procedure as the (asymptotic equivalence class of the) expected value of the KL divergence between the estimated and true distributions as a function of the number of samples. Using Fano method, we establish a fundamental limit on the information rate that can be achieved by any Bayesian estimator, and show that this fundamental limit is independent of the selection bias in the observational data. We characterize the Bayesian priors on the potential (factual and counterfactual) outcomes that achieve the optimal information rate. As a consequence, we show that a particular class of priors that have been widely used in the causal inference literature cannot achieve the optimal information rate. On the other hand, a broader class of priors can achieve the optimal information rate. We go on to propose a prior adaptation procedure (which we call the information based empirical Bayes procedure) that optimizes the Bayesian prior by maximizing an information theoretic criterion on the recovered causal effects rather than maximizing the marginal likelihood of the observed (factual) data. Building on our analysis, we construct an information optimal Bayesian causal inference algorithm.
References in corpus (10)
- Recursive Partitioning for Heterogeneous Causal Effects
- Rates of contraction of posterior distributions based on Gaussian process priors
- GPflow: A Gaussian process library using TensorFlow
- Reproducing kernel Hilbert spaces of Gaussian priors
- Lower bounds for posterior rates with Gaussian process priors
- Deep Counterfactual Networks with Propensity-Dropout
- Bandwidth selection for kernel density estimation with length-biased data
- Optimal Bayesian estimation in random covariate design with a rescaled Gaussian process prior
- Bayesian Regression Tree Ensembles that Adapt to Smoothness and Sparsity
- Model Criticism for Bayesian Causal Inference
Cited by in corpus (6)
- Nonparametric Estimation of Heterogeneous Treatment Effects: From Theory to Learning Algorithms
- Counterfactual Normalization: Proactively Addressing Dataset Shift and Improving Reliability Using Causal Mechanisms
- Causal-BALD: Deep Bayesian Active Learning of Outcomes to Infer Treatment-Effects from Observational Data
- Conditional Distributional Treatment Effect with Kernel Conditional Mean Embeddings and U-Statistic Regression
- Estimating Structural Target Functions using Machine Learning and Influence Functions
- Causal Link Discovery with Unequal Edge Error Tolerance