Generalization Error Curves for Analytic Spectral Algorithms under Power-law Decay
arXiv:2401.01599 · doi:10.1016/j.acha.2026.101920
Abstract
The generalization error curve of certain kernel regression method aims at determining the exact order of generalization error with various source condition, noise level and choice of the regularization parameter rather than the minimax rate. In this work, under mild assumptions, we rigorously provide a full characterization of the generalization error curves of the kernel gradient descent method (and a large class of analytic spectral algorithms) in kernel regression. Consequently, we could sharpen the near inconsistency of kernel interpolation and clarify the saturation effects of kernel regression algorithms with higher qualification, etc. Thanks to the neural tangent kernel theory, these results greatly improve our understanding of the generalization behavior of training the wide neural networks. A novel technical contribution, the analytic functional argument, might be of independent interest.
References in corpus (16)
- Neural Tangent Kernel: Convergence and Generalization in Neural Networks
- A Convergence Theory for Deep Learning via Over-Parameterization
- Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient Descent
- Boosting with early stopping: Convergence and consistency
- Just Interpolate: Kernel "Ridgeless" Regression Can Generalize
- Optimal Rates for Spectral Algorithms with Least-Squares Regression over Hilbert Spaces
- Kernel regression in high dimensions: Refined analysis beyond double descent
- Generalization Error Rates in Kernel Regression: The Crossover from the Noiseless to Noisy Regime
- Consistency of Interpolation with Laplace Kernels is a High-Dimensional Phenomenon
- Kernel interpolation generalizes poorly
- The Three Stages of Learning Dynamics in High-Dimensional Kernel Methods
- On the Asymptotic Learning Curves of Kernel Ridge Regression under Power-law Decay
- On the Saturation Effect of Kernel Ridge Regression
- On the Optimality of Misspecified Spectral Algorithms
- On the Optimality of Misspecified Kernel Ridge Regression
- Optimal Rate of Kernel Regression in Large Dimensions