Learning Over-Parametrized Two-Layer ReLU Neural Networks beyond NTK
arXiv:2007.04596
Abstract
We consider the dynamic of gradient descent for learning a two-layer neural network. We assume the input is drawn from a Gaussian distribution and the label of satisfies , where is a nonnegative vector and is an orthonormal matrix. We show that an over-parametrized two-layer neural network with ReLU activation, trained by gradient descent from random initialization, can provably learn the ground truth network with population loss at most in polynomial time with polynomial samples. On the other hand, we prove that any kernel method, including Neural Tangent Kernel, with a polynomial number of samples in , has population loss at least .
Conference on Learning Theory (COLT) 2020
References in corpus (10)
- Regularizing and Optimizing LSTM Language Models
- Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Optimization for deep learning: theory and algorithms
- Recovery Guarantees for One-hidden-layer Neural Networks
- Learning One-hidden-layer Neural Networks with Landscape Design
- Enhanced Convolutional Neural Tangent Kernels
- First-order Methods Almost Always Avoid Saddle Points
- Theoretical properties of the global optimizer of two layer neural network
- Shape Matters: Understanding the Implicit Bias of the Noise Covariance