Double/Debiased Machine Learning for Treatment and Causal Parameters
arXiv:1608.00060
Abstract
Most modern supervised statistical/machine learning (ML) methods are explicitly designed to solve prediction problems very well. Achieving this goal does not imply that these methods automatically deliver good estimators of causal parameters. Examples of such parameters include individual regression coefficients, average treatment effects, average lifts, and demand or supply elasticities. In fact, estimates of such causal parameters obtained via naively plugging ML estimators into estimating equations for such parameters can behave very poorly due to the regularization bias. Fortunately, this regularization bias can be removed by solving auxiliary prediction problems via ML tools. Specifically, we can form an orthogonal score for the target low-dimensional parameter by combining auxiliary and main ML predictions. The score is then used to build a de-biased estimator of the target parameter which typically will converge at the fastest possible 1/root(n) rate and be approximately unbiased and normal, and from which valid confidence intervals for these parameters of interest may be constructed. The resulting method thus could be called a "double ML" method because it relies on estimating primary and auxiliary predictive models. In order to avoid overfitting, our construction also makes use of the K-fold sample splitting, which we call cross-fitting. This allows us to use a very broad set of ML predictive methods in solving the auxiliary and main prediction problems, such as random forest, lasso, ridge, deep neural nets, boosted trees, as well as various hybrids and aggregators of these methods.
71 pages, 2 figures
References in corpus (11)
- Simultaneous analysis of Lasso and Dantzig selector
- Confidence Intervals and Hypothesis Testing for High-Dimensional Regression
- Least squares after model selection in high-dimensional sparse models
- Square-Root Lasso: Pivotal Recovery of Sparse Signals via Conic Programming
- Robust Inference on Average Treatment Effects with Possibly More Covariates than Observations
- Sparse Estimators and the Oracle Property, or the Return of Hodges' Estimator
- Adaptive Concentration of Regression Trees, with Application to Random Forests
- High-dimensional instrumental variables regression and confidence sets
- Inference for High-Dimensional Sparse Econometric Models
- LASSO Methods for Gaussian Instrumental Variables Models
- On the Distribution of Penalized Maximum Likelihood Estimators: The LASSO, SCAD, and Thresholding
Cited by in corpus (27)
- Matrix Completion Methods for Causal Panel Data Models
- Estimating individual treatment effect: generalization bounds and algorithms
- Bayesian Inference of Individualized Treatment Effects using Multi-task Gaussian Processes
- Causality Learning: A New Perspective for Interpretable Machine Learning
- Learning Weighted Representations for Generalization Across Designs
- Counterfactual Prediction with Deep Instrumental Variables Networks
- Confounding-Robust Policy Improvement
- DeepMatch: Balancing Deep Covariate Representations for Causal Inference Using Adversarial Training
- On the multiply robust estimation of the mean of the g-functional
- Robust Estimation of Causal Effects via High-Dimensional Covariate Balancing Propensity Score
- Semiparametric Contextual Bandits
- Causal Inference for Comprehensive Cohort Studies
- Robust Inference for Mediated Effects in Partially Linear Models
- Learning When-to-Treat Policies
- A Deep Causal Inference Approach to Measuring the Effects of Forming Group Loans in Online Non-profit Microfinance Platform
- Policy Evaluation with Latent Confounders via Optimal Balance
- Faster Rates for Policy Learning
- Causal Inference: A Missing Data Perspective
- Semiparametric panel data models using neural networks
- Data-adaptive doubly robust instrumental variable methods for treatment effect heterogeneity
- Generalised linear models for prognosis and intervention: Theory, practice, and implications for machine learning
- Single Point Transductive Prediction
- Estimating Average Treatment Effects: Supplementary Analyses and Remaining Challenges
- Nonparametric causal effects based on incremental propensity score interventions
- Statistical Inference for Data-adaptive Doubly Robust Estimators with Survival Outcomes
- Pricing Engine: Estimating Causal Impacts in Real World Business Settings
- Off-Policy Evaluation via Adaptive Weighting with Data from Contextual Bandits