Optimal Rates for Averaged Stochastic Gradient Descent under Neural Tangent Kernel Regime
arXiv:2006.12297
Abstract
We analyze the convergence of the averaged stochastic gradient descent for overparameterized two-layer neural networks for regression problems. It was recently found that a neural tangent kernel (NTK) plays an important role in showing the global convergence of gradient-based methods under the NTK regime, where the learning dynamics for overparameterized neural networks can be almost characterized by that for the associated reproducing kernel Hilbert space (RKHS). However, there is still room for a convergence rate analysis in the NTK regime. In this study, we show that the averaged stochastic gradient descent can achieve the minimax optimal convergence rate, with the global convergence guarantee, by exploiting the complexities of the target function and the RKHS associated with the NTK. Moreover, we show that the target function specified by the NTK of a ReLU network can be learned at the optimal convergence rate through a smooth approximation of a ReLU network under certain conditions.
35 pages
References in corpus (3)
Cited by in corpus (7)
- How Neural Networks Extrapolate: From Feedforward to Graph Neural Networks
- Regularization Matters: A Nonparametric Perspective on Overparametrized Neural Network
- Universal scaling laws in the gradient descent training of neural networks
- On the Double Descent of Random Features Models Trained with SGD
- Learning curves for Gaussian process regression with power-law priors and targets
- Painless step size adaptation for SGD
- Neural Optimization Kernel: Towards Robust Deep Learning