Coarsening Latent-Class Probabilities: Directional Distortion and Coverage Loss
arXiv:2608.11784
Abstract
Outcomes are regressed on a calibrated probability vector for unobserved class membership. Under a structural conditional mean excluding the score and conditional calibration, the observed-data model reduces to a partially linear regression. The probability vector is a Berkson-type surrogate for membership, so the effect vector is identified without attenuation. In practice the vector is often coarsened to a hard label - an argmax, a confidence threshold - and that label need not retain the Berkson property. For any coarsening the plug-in estimator converges to , where is determined by the regression of the discarded part of the score on the retained part. Coarsening therefore leaves undistorted exactly when that regression vanishes, and otherwise distorts some contrasts far more than others. The same operator determines the bias that drives coverage loss. Where that bias is of the order of the standard error, the Wald interval has limiting coverage , with their ratio. A fixed bias sends coverage to zero. Operator, standard error, and - through the uncoarsened estimator - the bias are estimable from observed data, so the implied coverage can be approximated before the interval is reported. Simulations show severe coverage loss after argmax coarsening. Three real-data audits exhibit the direction-specific distortion.
42 pages, 6 figures, 11 tables. Includes supplementary material with all proofs, two further identification results, additional experiments, and a land-cover audit