Sampling techniques for big data analysis in finite population inference
arXiv:1801.09728 · doi:10.1111/insr.12290
Abstract
In analyzing big data for finite population inference, it is critical to adjust for the selection bias in the big data. In this paper, we propose two methods of reducing the selection bias associated with the big data sample. The first method uses a version of inverse sampling by incorporating auxiliary information from external sources, and the second one borrows the idea of data integration by combining the big data sample with an independent probability sample. Two simulation studies show that the proposed methods are unbiased and have better coverage rates than their alternatives. In addition, the proposed methods are easy to implement in practice.
24 pages, 3 tables
References in corpus (1)
Cited by in corpus (4)
- Estimation of the size of informal employment based on administrative records with non-ignorable selection mechanism
- Reconstructing Curves from Sparse Samples on Riemannian Manifolds
- Enhancing the Demand for Labour survey by including skills from online job advertisements using model-assisted calibration
- Double Robust Mass-Imputation with Matching Estimators