An Analytical Formula of Population Gradient for two-layered ReLU network and its Applications in Convergence and Critical Point Analysis
arXiv:1703.00560
Abstract
In this paper, we explore theoretical properties of training a two-layered ReLU network with centered -dimensional spherical Gaussian input (=ReLU). We train our network with gradient descent on to mimic the output of a teacher network with the same architecture and fixed parameters . We show that its population gradient has an analytical formula, leading to interesting theoretical analysis of critical points and convergence behaviors. First, we prove that critical points outside the hyperplane spanned by the teacher parameters ("out-of-plane") are not isolated and form manifolds, and characterize in-plane critical-point-free regions for two ReLU case. On the other hand, convergence to for one ReLU node is guaranteed with at least probability, if weights are initialized randomly with standard deviation upper-bounded by , consistent with empirical practice. For network with many ReLU nodes, we prove that an infinitesimal perturbation of weight initialization results in convergence towards (or its permutation), a phenomenon known as spontaneous symmetric-breaking (SSB) in physics. We assume no independence of ReLU activations. Simulation verifies our findings.
International Conference on Machine Learning (ICML) 2017
References in corpus (4)
Cited by in corpus (48)
- Gradient Descent Provably Optimizes Over-parameterized Neural Networks
- Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Learning and Generalization in Overparameterized Neural Networks, Going Beyond Two Layers
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning
- Learning One-hidden-layer Neural Networks with Landscape Design
- A Convergence Analysis of Gradient Descent for Deep Linear Neural Networks
- Implicit Full Waveform Inversion with Deep Neural Representation
- What Can ResNet Learn Efficiently, Going Beyond Kernels?
- Understanding Straight-Through Estimator in Training Activation Quantized Neural Nets
- Towards Theoretically Understanding Why SGD Generalizes Better Than ADAM in Deep Learning
- On the loss landscape of a class of deep neural networks with no bad local valleys
- Generalization Error Bounds of Gradient Descent for Learning Over-parameterized Deep ReLU Networks
- Towards a Mathematical Understanding of Neural Network-Based Machine Learning: what we know and what we don't
- An Improved Analysis of Training Over-parameterized Deep Neural Networks
- SGD Learns Over-parameterized Networks that Provably Generalize on Linearly Separable Data
- Understanding self-supervised Learning Dynamics without Contrastive Pairs
- Width Provably Matters in Optimization for Deep Linear Neural Networks
- Critical Points of Neural Networks: Analytical Forms and Landscape Properties
- A Generalized Neural Tangent Kernel Analysis for Two-layer Neural Networks
- Theory-Inspired Path-Regularized Differential Network Architecture Search
- Porcupine Neural Networks: (Almost) All Local Optima are Global
- Algorithmic Regularization in Learning Deep Homogeneous Models: Layers are Automatically Balanced
- A Provably Correct Algorithm for Deep Learning that Actually Works
- Blended Coarse Gradient Descent for Full Quantization of Deep Neural Networks
- Can SGD Learn Recurrent Neural Networks with Provable Generalization?
- Learning Two Layer Rectified Neural Networks in Polynomial Time
- The Multilinear Structure of ReLU Networks
- Learning a Single Neuron with Gradient Methods
- Hardness of Learning Neural Networks with Natural Weights
- On Connected Sublevel Sets in Deep Learning
- Learning Over-Parametrized Two-Layer ReLU Neural Networks beyond NTK
- Improved Learning of One-hidden-layer Convolutional Neural Networks with Overlaps
- Convergence of a Relaxed Variable Splitting Method for Learning Sparse Neural Networks via , and transformed- Penalties
- The Effects of Mild Over-parameterization on the Optimization Landscape of Shallow ReLU Neural Networks
- The Global Optimization Geometry of Shallow Linear Neural Networks
- Sub-Optimal Local Minima Exist for Neural Networks with Almost All Non-Linear Activations
- Forward Super-Resolution: How Can GANs Learn Hierarchical Generative Models for Real-World Distributions
- Nonlinear Inductive Matrix Completion based on One-layer Neural Networks
- Collective evolution of weights in wide neural networks
- From Local Pseudorandom Generators to Hardness of Learning
- Naive Gabor Networks for Hyperspectral Image Classification
- A Unified Framework for Training Neural Networks
- Making Method of Moments Great Again? -- How can GANs learn distributions
- Nonparametric Learning of Two-Layer ReLU Residual Units
- A Modular Analysis of Provable Acceleration via Polyak's Momentum: Training a Wide ReLU Network and a Deep Linear Network
- Convergence of a Relaxed Variable Splitting Coarse Gradient Descent Method for Learning Sparse Weight Binarized Activation Neural Networks
- Understanding Modern Techniques in Optimization: Frank-Wolfe, Nesterov's Momentum, and Polyak's Momentum