paper

Semiparametric Efficient Data Integration Using the Dual-Frame Sampling Framework

arXiv:2601.08707

Abstract

Integrating probability and non-probability samples is increasingly important, yet unknown non-probability inclusion mechanisms complicate identification and efficient estimation. We develop semiparametric theory for dual-frame data integration and propose two complementary estimators. The first models the conditional inclusion probability parametrically and, under correct specification, attains the semiparametric efficiency bound. The probability sample serves as an outcome-observed validation sample, identifying the conditional inclusion parameters without instrumental variables even under informative selection. We derive the efficient influence function while leaving the conditional distribution of the design inclusion probability unrestricted. The resulting estimating equations satisfy exact augmentation invariance, which supports both same-sample sieve estimation and pooled cross-fitting under weak nuisance-rate conditions. The second estimator, motivated by a two-stage sampling approximation, does not require a non-probability inclusion model; although not fully efficient, it is efficient within a restricted augmentation class, is robust to misspecification of the non-probability inclusion model, and its oracle version weakly dominates optimally augmented probability-sample-only inference. Simulations and a public-data repeated-sampling study show efficiency gains under correct specification and stable performance under misspecification and weak identification.