Propensity score analysis with latent covariates: Measurement error bias correction using the covariate's posterior mean, aka the inclusive factor score
arXiv:1907.12709 · doi:10.3102/1076998620911920
Abstract
We address measurement error bias in propensity score (PS) analysis due to covariates that are latent variables. In the setting where latent covariate is measured via multiple error-prone items , PS analysis using several proxies for -- the items themselves, a summary score (mean/sum of the items), or the conventional factor score (cFS , i.e., predicted value of based on the measurement model) -- often results in biased estimation of the causal effect, because balancing the proxy (between exposure conditions) does not balance . We propose an improved proxy: the conditional mean of given the combination of , the observed covariates , and exposure , denoted . The theoretical support, which applies whether is latent or not (but is unobserved), is that balancing (e.g., via weighting or matching) implies balancing the mean of . For a latent , we estimate by the inclusive factor score (iFS) -- predicted value of from a structural equation model that captures the joint distribution of given . Simulation shows that PS analysis using the iFS substantially improves balance on the first five moments of and reduces bias in the estimated causal effect. Hence, within the proxy variables approach, we recommend this proxy over existing ones. We connect this proxy method to known results about weighting/matching functions (Lockwood & McCaffrey, 2016; McCaffrey, Lockwood, & Setodji, 2013). We illustrate the method in handling latent covariates when estimating the effect of out-of-school suspension on risk of later police arrests using Add Health data.