paper

Orthogonal double residual learning for optimal individualized treatment rules

arXiv:2608.24085

Abstract

Individualized treatment rules (ITRs) map baseline characteristics to treatment recommendations, with the optimal ITR maximizing expected reward or policy welfare. Indirect methods may require restrictive modeling assumptions, whereas direct methods can be sensitive to nuisance estimation error and limited overlap. We propose orthogonal double residual learning (ODRL), a two-stage, cross-fitted framework that directly targets the optimal ITR through cost-sensitive classification using the product of treatment and outcome residuals. To our knowledge, ODRL is the first direct method with a universally Neyman orthogonal objective requiring neither restrictive modeling assumptions nor inverse propensity score weighting. Thus, nuisance estimation errors affect regret through a second-order product, and ODRL remains robust under limited overlap. The Fisher consistent objective accommodates general decision rule sieves. We establish nonasymptotic high probability value function regret bounds relative to the Bayes classifier for VC classes, including linear rules and decision trees, and calibrated regret bounds for surrogate relaxations using support vector machines and deep ReLU neural networks. We further show that generic surrogate relaxations need not preserve orthogonality, whereas bounded score hinge learning does. Simulations demonstrate strong performance across complex and linear decision boundaries, limited overlap, and working model misspecification. Applications to the Right Heart Catheterization study and the Oxford Net Zero experiment illustrate interpretable treatment or policy recommendations. The \texttt{odrlITR} R package implements ODRL.

Orthogonal double residual learning for optimal individualized treatment rules · wovepaper