Detecting Where Effects Occur by Testing Hypotheses in Order
arXiv:2602.21068
Abstract
Experimental evaluations of public policies often randomize a new intervention within many sites or blocks. After an overall statistically significant result is reported, the natural question from a policy maker is: \emph{where} did effects occur? Standard adjustments for multiple testing answer this question with little power because they ignore how the experiment is organized: blocks nest within cohorts, sites, and districts. We organize the hypotheses in the shape of a tree that follows this administrative structure and test them top-down, stopping at any branch where the null is not rejected. A stopping rule and valid tests at each node suffice for weak control of the family-wise error rate (FWER). Whether the unadjusted procedure also controls the FWER in the strong sense depends on an \emph{error load} computable from design quantities before any data are tested; when the load exceeds one, an adaptive -schedule, which we prove controls the FWER on regular and irregular trees without pruning, restores control. [Correction, August 2026: an anonymous referee identified errors in the previous version. The error loads reported in the paper were computed incorrectly; corrected loads exceed 1 in all 25 block-randomized MDRC education trials at the planning effect size , so the adjustment the earlier version claimed unnecessary is in fact required. The theorem claiming FWER control under branch pruning and its switching corollary are false as stated and are withdrawn; an exact counterexample attains FWER 0.063 at . The detection comparison with the Hommel procedure is under recomputation. A correction notice on page 1 details what changes and what stands; a fully corrected version is in preparation.]