Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
arXiv:1901.08584
Abstract
Recent works have cast some light on the mystery of why deep nets fit any data and generalize despite being very overparametrized. This paper analyzes training and generalization for a simple 2-layer ReLU net with random initialization, and provides the following improvements over recent works: (i) Using a tighter characterization of training speed than recent papers, an explanation for why training a neural net with random labels leads to slower training, as originally observed in [Zhang et al. ICLR'17]. (ii) Generalization bound independent of network size, using a data-dependent complexity measure. Our measure distinguishes clearly between random labels and true labels on MNIST and CIFAR, as shown by experiments. Moreover, recent papers require sample complexity to increase (slowly) with the size, while our sample complexity is completely independent of the network size. (iii) Learnability of a broad class of smooth functions by 2-layer ReLU nets trained via gradient descent. The key idea is to track dynamics of training and generalization via properties of a related kernel.
In ICML 2019
References in corpus (5)
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Recovery Guarantees for One-hidden-layer Neural Networks
- Diverse Neural Network Learns True Target Functions
- Generalization Bounds of SGLD for Non-convex Learning: Two Theoretical Viewpoints
- Critical Points of Neural Networks: Analytical Forms and Landscape Properties
Cited by in corpus (35)
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains
- Proving the Lottery Ticket Hypothesis: Pruning is All You Need
- Environmental drivers of systematicity and generalization in a situated agent
- An Improved Analysis of Training Over-parameterized Deep Neural Networks
- Asymptotics of Wide Networks from Feynman Diagrams
- A Selective Overview of Deep Learning
- Finite Versus Infinite Neural Networks: an Empirical Study
- Generalization Guarantees for Neural Networks via Harnessing the Low-rank Structure of the Jacobian
- Gradient Dynamics of Shallow Univariate ReLU Networks
- Explicitizing an Implicit Bias of the Frequency Principle in Two-layer Neural Networks
- Distillation Early Stopping? Harvesting Dark Knowledge Utilizing Anisotropic Information Retrieval For Overparameterized Neural Network
- Shape Matters: Understanding the Implicit Bias of the Noise Covariance
- Neural Networks Learning and Memorization with (almost) no Over-Parameterization
- Analysis of the Gradient Descent Algorithm for a Deep Neural Network Model with Skip-connections
- Asymptotics of Wide Convolutional Neural Networks
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural Networks
- The Surprising Simplicity of the Early-Time Learning Dynamics of Neural Networks
- Optimization Theory for ReLU Neural Networks Trained with Normalization Layers
- Harnessing the Power of Infinitely Wide Deep Nets on Small-data Tasks
- WeMix: How to Better Utilize Data Augmentation
- Decoupling Gating from Linearity
- Over Parameterized Two-level Neural Networks Can Learn Near Optimal Feature Representations
- Neural tangent kernels, transportation mappings, and universal approximation
- Disentangling Adaptive Gradient Methods from Learning Rates
- Shallow Univariate ReLu Networks as Splines: Initialization, Loss Surface, Hessian, & Gradient Flow Dynamics
- Learning Over-Parametrized Two-Layer ReLU Neural Networks beyond NTK
- Exploring Weight Importance and Hessian Bias in Model Pruning
- Generalization Error of Generalized Linear Models in High Dimensions
- Learning Boolean Circuits with Neural Networks
- On the Learning Dynamics of Two-layer Nonlinear Convolutional Neural Networks
- Deep Gated Networks: A framework to understand training and generalisation in deep learning
- Analytic expressions for the output evolution of a deep neural network
- Learning the gravitational force law and other analytic functions
- On Symmetry and Initialization for Neural Networks
- Nearly Minimal Over-Parametrization of Shallow Neural Networks