When Does -Boosting Overfit Benignly? High-Dimensional Risk Asymptotics and the Implicit Bias
arXiv:2605.06314
Abstract
Benign overfitting is well-characterized in geometries, but its behavior under the implicit bias of greedy ensembles remains challenging. The analytical barrier stems from the non-linear coupling of coordinate selection thresholds, which invalidates standard spectral resolvent tools. To isolate this algorithmic bias, we characterize the high-dimensional risk of continuous-time -Boosting over features and samples. By coupling the Convex Gaussian Minimax Theorem with delicate asymptotic expansions of double-sided truncated Gaussian moments, we analytically resolve the non-smooth interpolant. Under an isotropic pure-noise model, we prove that benign overfitting fails at the linear rate: greedy selection localizes noise into sparse active sets, and the excess variance decays at a logarithmic rate for noise variance . We remark that while this localization mechanism should persist in the presence of signals, the exact signal-noise decomposition remains an open problem. For spiked-isotropic designs with head eigenvalues and tail dimensions, the risk converges to zero when , but only at a logarithmic rate , which is slower than the linear decay observed in geometries. To avoid this slow convergence, we analyze the non-smooth subdifferential dynamics of the boosting flow. This yields a tuning-free early stopping rule that, under a bounded -path condition, recovers the Lasso basic inequality and attains the minimax-optimal empirical prediction rate for -bounded signals.