Statistical inference for the mean outcome under a possibly non-unique optimal treatment strategy
arXiv:1603.07573 · doi:10.1214/15-AOS1384
Abstract
We consider challenges that arise in the estimation of the mean outcome under an optimal individualized treatment strategy defined as the treatment rule that maximizes the population mean outcome, where the candidate treatment rules are restricted to depend on baseline covariates. We prove a necessary and sufficient condition for the pathwise differentiability of the optimal value, a key condition needed to develop a regular and asymptotically linear (RAL) estimator of the optimal value. The stated condition is slightly more general than the previous condition implied in the literature. We then describe an approach to obtain root- rate confidence intervals for the optimal value even when the parameter is not pathwise differentiable. We provide conditions under which our estimator is RAL and asymptotically efficient when the mean outcome is pathwise differentiable. We also outline an extension of our approach to a multiple time point problem. All of our results are supported by simulations.
Published at http://dx.doi.org/10.1214/15-AOS1384 in the Annals of Statistics (http://www.imstat.org/aos/) by the Institute of Mathematical Statistics (http://www.imstat.org)
References in corpus (4)
Cited by in corpus (36)
- Statistical inference for the mean outcome under a possibly non-unique optimal treatment strategy
- Quasi-Oracle Estimation of Heterogeneous Treatment Effects
- Sensitivity analysis via the proportion of unmeasured confounding
- Modern approaches for evaluating treatment effect heterogeneity from clinical trials and observational data
- Experimental Evaluation of Individualized Treatment Rules
- Bridging the gap: Towards an Expanded Toolkit for AI-driven Decision-Making in the Public Sector
- Statistical Inference for Online Decision Making via Stochastic Gradient Descent
- Statistical Inference of the Value Function for Reinforcement Learning in Infinite Horizon Settings
- Inference for Batched Bandits
- Estimating heterogeneous treatment effects with right-censored data via causal survival forests
- Post-Contextual-Bandit Inference
- Treatment Effects on Ordinal Outcomes: Causal Estimands and Sharp Bounds
- The Optimal Dynamic Treatment Rule SuperLearner: Considerations, Performance, and Application
- Median Optimal Treatment Regimes
- Off-Policy Evaluation of Bandit Algorithm from Dependent Samples under Batch Update Policy
- Causal Inference with Corrupted Data: Measurement Error, Missing Values, Discretization, and Differential Privacy
- Adaptive Sequential Design for a Single Time-Series
- Doubly Robust Interval Estimation for Optimal Policy Evaluation in Online Learning
- Jump Interval-Learning for Individualized Decision Making
- Estimation and inference on high-dimensional individualized treatment rule in observational data using split-and-pooled de-correlated score
- Pessimistic Model Selection for Offline Deep Reinforcement Learning
- Enhanced inference for distributions and quantiles of individual treatment effects in various experiments
- Resampling-based Confidence Intervals for Model-free Robust Inference on Optimal Treatment Regimes
- Dirac Delta Regression: Conditional Density Estimation with Clinical Trials
- Positivity-free Policy Learning with Observational Data
- Calibrated Optimal Decision Making with Multiple Data Sources and Limited Outcome
- Adaptive Doubly Robust Estimator from Non-stationary Logging Policy under a Convergence of Average Probability
- Model-Assisted Uniformly Honest Inference for Optimal Treatment Regimes in High Dimension
- A framework for causal segmentation analysis with machine learning in large-scale digital experiments
- Finding the Optimal Dynamic Treatment Regime Using Smooth Fisher Consistent Surrogate Loss
- Online Testing of Subgroup Treatment Effects Based on Value Difference
- Performance and Application of Estimators for the Value of an Optimal Dynamic Treatment Rule
- A Comprehensive Framework for the Evaluation of Individual Treatment Rules From Observational Data
- Off-Policy Evaluation via Adaptive Weighting with Data from Contextual Bandits
- A Statistical Test for the Benefits of Personalizing Interventions
- The Adaptive Doubly Robust Estimator for Policy Evaluation in Adaptive Experiments and a Paradox Concerning Logging Policy