High-dimensional semi-supervised learning: in search for optimal inference of the mean
arXiv:1902.00772 · doi:10.1093/biomet/asab042
Abstract
A fundamental challenge in semi-supervised learning lies in the observed data's disproportional size when compared with the size of the data collected with missing outcomes. An implicit understanding is that the dataset with missing outcomes, being significantly larger, ought to improve estimation and inference. However, it is unclear to what extent this is correct. We illustrate one clear benefit: root-n inference of the outcome's mean is possible while only requiring a consistent estimation of the outcome, possibly at a rate slower than root n. This is achieved by a novel k-fold, cross-fitted, double robust estimator. We discuss both linear and nonlinear outcomes. Such an estimator is particularly suited for models that naturally do not admit root-n consistency, such as high-dimensional, nonparametric or semiparametric models. We apply our methods to estimating heterogeneous treatment effects.
References in corpus (11)
- Meta-learners for Estimating Heterogeneous Treatment Effects using Machine Learning
- On asymptotically optimal confidence regions and tests for high-dimensional models
- On the conditions used to prove oracle results for the Lasso
- Square-Root Lasso: Pivotal Recovery of Sparse Signals via Conic Programming
- SLOPE - Adaptive variable selection via convex optimization
- Efficient and Adaptive Linear Regression in Semi-Supervised Settings
- High-dimensional inference in misspecified linear models
- Asymptotic behavior of -based Laplacian regularization in semi-supervised learning
- Bootstrapping and Sample Splitting For High-Dimensional, Assumption-Free Inference
- Minimax-optimal semi-supervised regression on unknown manifolds
- Semi-Supervised linear regression