Rerandomization for covariate balance mitigates -hacking caused by strategically selecting covariates in regression adjustment
arXiv:2505.01137
Abstract
Rerandomization enforces covariate balance across treatment groups in the design stage of experiments. Despite its intuitive appeal, its theoretical justification remains unsatisfying because its benefits of improving efficiency for estimating the average treatment effect diminish if we use regression adjustment in the analysis stage. To strengthen the theory of rerandomization, we show that it mitigates false discoveries resulting from -hacking caused by strategically selecting covariates in regression adjustment to get more significant -values. Moreover, we show that rerandomization with a sufficiently stringent threshold can resolve such -hacking. As a byproduct, our theory offers guidance for choosing the threshold in rerandomization in practice.
67 pages, 5 figures. Revised manuscript and supplementary material; theoretical discussion and simulation studies expanded