Regularization Matters: Generalization and Optimization of Neural Nets v.s. their Induced Kernel
arXiv:1810.05369
Abstract
Recent works have shown that on sufficiently over-parametrized neural nets, gradient descent with relatively large initialization optimizes a prediction function in the RKHS of the Neural Tangent Kernel (NTK). This analysis leads to global convergence results but does not work when there is a standard regularizer, which is useful to have in practice. We show that sample efficiency can indeed depend on the presence of the regularizer: we construct a simple distribution in d dimensions which the optimal regularized neural net learns with samples but the NTK requires samples to learn. To prove this, we establish two analysis tools: i) for multi-layer feedforward ReLU nets, we show that the global minimizer of a weakly-regularized cross-entropy loss is the max normalized margin solution among all neural nets, which generalizes well; ii) we develop a new technique for proving lower bounds for kernel methods, which relies on showing that the kernel cannot focus on informative features. Motivated by our generalization results, we study whether the regularized global optimum is attainable. We prove that for infinite-width two-layer nets, noisy gradient descent optimizes the regularized neural net loss to a global minimum in polynomial iterations.
version 2: title changed from originally "On the Margin Theory of Feedforward Neural Networks". Substantial changes from old version of paper, including a new lower bound on NTK sample complexity version 3: reorganized NTK lower bound proof version 4: reorganized proof of optimization result
Cited by in corpus (25)
- Learning Imbalanced Datasets with Label-Distribution-Aware Margin Loss
- Generalization Bounds of Stochastic Gradient Descent for Wide and Deep Neural Networks
- The Role of Neural Network Activation Functions
- Gradient Descent Maximizes the Margin of Homogeneous Neural Networks
- What Can ResNet Learn Efficiently, Going Beyond Kernels?
- Generalization Error Bounds of Gradient Descent for Learning Over-parameterized Deep ReLU Networks
- Classifying high-dimensional Gaussian mixtures: Where kernel methods fail and neural networks succeed
- A Generalized Neural Tangent Kernel Analysis for Two-layer Neural Networks
- How Much Over-parameterization Is Sufficient to Learn Deep ReLU Networks?
- Mathematical Models of Overparameterized Neural Networks
- Implicit regularization for deep neural networks driven by an Ornstein-Uhlenbeck like process
- Birth-death dynamics for sampling: Global convergence, approximations and their asymptotics
- Generalisation Guarantees for Continual Learning with Orthogonal Gradient Descent
- Computationally Efficient Feature Significance and Importance for Machine Learning Models
- Inductive Bias of Multi-Channel Linear Convolutional Networks with Bounded Weight Norm
- Global Convergence of Three-layer Neural Networks in the Mean Field Regime
- Temporal-difference learning with nonlinear function approximation: lazy training and mean field regimes
- On the Global Convergence of Gradient Descent for multi-layer ResNets in the mean-field regime
- Global Convergence of SGD On Two Layer Neural Nets
- Properties of the After Kernel
- A Modular Analysis of Provable Acceleration via Polyak's Momentum: Training a Wide ReLU Network and a Deep Linear Network
- Understanding Modern Techniques in Optimization: Frank-Wolfe, Nesterov's Momentum, and Polyak's Momentum
- A note on regularised NTK dynamics with an application to PAC-Bayesian training
- Recent Advances in Large Margin Learning
- Provable Generalization of SGD-trained Neural Networks of Any Width in the Presence of Adversarial Label Noise