Symmetry & critical points for a model shallow neural network
arXiv:2003.10576 · doi:10.1016/j.physd.2021.133014
Abstract
We consider the optimization problem associated with fitting two-layer ReLU networks with hidden neurons, where labels are assumed to be generated by a (teacher) neural network. We leverage the rich symmetry exhibited by such models to identify various families of critical points and express them as power series in . These expressions are then used to derive estimates for several related quantities which imply that not all spurious minima are alike. In particular, we show that while the loss function at certain types of spurious minima decays to zero like , in other cases the loss converges to a strictly positive constant. The methods used depend on symmetry, the geometry of group actions, bifurcation, and Artin's implicit function theorem.
References in corpus (13)
- Deep Learning in Neural Networks: An Overview
- Searching for Activation Functions
- Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
- Qualitatively characterizing neural network optimization problems
- Recovery Guarantees for One-hidden-layer Neural Networks
- Modelling the influence of data structure on learning in neural networks: the hidden manifold model
- Globally Optimal Gradient Descent for a ConvNet with Gaussian Inputs
- Diverse Neural Network Learns True Target Functions
- Local minima in training of neural networks
- Porcupine Neural Networks: (Almost) All Local Optima are Global
- Asynchronous Networks: Modularization of Dynamics Theorem
- Analytic Characterization of the Hessian in Shallow ReLU Models: A Tale of Symmetry
- Symmetry Breaking in Symmetric Tensor Decomposition
Cited by in corpus (4)
- Dynamic Neural Diversification: Path to Computationally Sustainable Neural Networks
- On the Stability Properties and the Optimization Landscape of Training Problems with Squared Loss for Neural Networks and General Nonlinear Conic Approximation Schemes
- Analytic Study of Families of Spurious Minima in Two-Layer ReLU Neural Networks: A Tale of Symmetry II
- Equivariant bifurcation, quadratic equivariants, and symmetry breaking for the standard representation of