Counterfactual Optimization of Policy Interventions: Lexical Ordering and Leapfrogging
arXiv:2608.20505
Abstract
Most data-driven policy learning methods maximize average outcomes, overlooking the possibility that a policy beneficial on average may still harm a substantial fraction of individuals. Motivated by the ethical principle of "first do no harm", we study how to design a change from a baseline policy that improves overall welfare while keeping the worst-case probability or expectation of individual harm below a specified limit. We establish sufficient conditions under which an optimal policy transition has a lexical leapfrogging structure: groups defined by covariates and current treatment are ranked by a priority score, and any treatment change moves them directly to the conditionally optimal treatment. We derive this score under several models for the dependence among potential outcomes. We demonstrate this harm-aware policy optimization approach in a reanalysis of the I-SPY2 breast cancer platform trial and show how the consideration of counterfactual harm may lead to different conclusions about which treatment-subgroup pairs may warrant deprioritization in further clinical evaluation.