paper

Debiased Machine Learning with Many Cross-Fitting Folds

arXiv:2411.01864

Abstract

This paper studies debiased machine learning (DML) when the number of cross-fitting folds, , may grow with the sample size . Existing fixed- asymptotic theory implies that DML1 and DML2, the two main DML variants, are asymptotically equivalent, providing no guidance on which variant to use or how to choose . We show that this equivalence can break down when grows proportionally to : DML1 can exhibit asymptotic bias, in which case standard inference based on DML1 fails---as can occur, for instance, for the local average treatment effect (LATE)---whereas inference based on DML2 remains valid. Moreover, we show that, under an algorithmic-stability condition, estimation and inference based on DML2 are valid for any , including the leave-one-out case, . Finally, for scalar DML2 estimators whose first-step estimators admit a stochastic linear expansion, we derive a second-order approximation showing that larger values of reduce the second-order asymptotic bias and mean-squared error, although the marginal improvements diminish.

This paper was previously circulated under the title "On the Asymptotic Properties of Debiased Machine Learning Estimators.'' Now it has 70 pages and 6 figures

Debiased Machine Learning with Many Cross-Fitting Folds · wovepaper