Doing Great at Estimating CATE? On the Neglected Assumptions in Benchmark Comparisons of Treatment Effect Estimators
arXiv:2107.13346
Abstract
The machine learning toolbox for estimation of heterogeneous treatment effects from observational data is expanding rapidly, yet many of its algorithms have been evaluated only on a very limited set of semi-synthetic benchmark datasets. In this paper, we show that even in arguably the simplest setting -- estimation under ignorability assumptions -- the results of such empirical evaluations can be misleading if (i) the assumptions underlying the data-generating mechanisms in benchmark datasets and (ii) their interplay with baseline algorithms are inadequately discussed. We consider two popular machine learning benchmark datasets for evaluation of heterogeneous treatment effect estimators -- the IHDP and ACIC2016 datasets -- in detail. We identify problems with their current use and highlight that the inherent characteristics of the benchmark datasets favor some algorithms over others -- a fact that is rarely acknowledged but of immense relevance for interpretation of empirical results. We close by discussing implications and possible next steps.
Workshop on the Neglected Assumptions in Causal Inference at the International Conference on Machine Learning (ICML), 2021
References in corpus (6)
- Discussion on "Bayesian Regression Tree Models for Causal Inference: Regularization, Confounding, and Heterogeneous Effects" by Hahn, Murray and Carvalho
- Bayesian Inference of Individualized Treatment Effects using Multi-task Gaussian Processes
- Nonparametric Estimation of Heterogeneous Treatment Effects: From Theory to Learning Algorithms
- Atlantic Causal Inference Conference (ACIC) Data Analysis Challenge 2017
- Learning Overlapping Representations for the Estimation of Individualized Treatment Effects
- RealCause: Realistic Causal Inference Benchmarking