Localising Dropout Variance in Twin Networks
arXiv:2507.03622
Abstract
Accurate individual treatment-effect estimation demands not only reliable point predictions but also uncertainty measures that help practitioners \emph{locate} the source of model failure. We introduce a layer-wise variance decomposition for deep twin-network models: by toggling Monte Carlo Dropout independently in the shared encoder and the outcome heads, we split total predictive variance into an \emph{encoder component} () and a \emph{head component} (), with by the law of total variance. Across three synthetic covariate-shift regimes, the encoder component dominates under distributional shift () while the head component becomes informative only once encoder uncertainty is controlled. On a real-world twins cohort with induced multivariate shift, only spikes on out-of-distribution samples and becomes the primary error predictor (), while remains flat. The decomposition adds negligible cost over standard MC Dropout and provides a practical diagnostic for deciding whether to collect more diverse covariates or more outcome data.
14 pages, 5 figures, 3 tables