Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances
arXiv:2105.12221
Abstract
We study how permutation symmetries in overparameterized multi-layer neural networks generate `symmetry-induced' critical points. Assuming a network with layers of minimal widths reaches a zero-loss minimum at isolated points that are permutations of one another, we show that adding one extra neuron to each layer is sufficient to connect all these previously discrete minima into a single manifold. For a two-layer overparameterized network of width we explicitly describe the manifold of global minima: it consists of affine subspaces of dimension at least that are connected to one another. For a network of width , we identify the number of affine subspaces containing only symmetry-induced critical points that are related to the critical points of a smaller network of width . Via a combinatorial analysis, we derive closed-form formulas for and and show that the number of symmetry-induced critical subspaces dominates the number of affine subspaces forming the global minima manifold in the mildly overparameterized regime (small ) and vice versa in the vastly overparameterized regime (). Our results provide new insights into the minimization of the non-convex loss function of overparameterized neural networks.
29 pages, 12 figures, ICML 2021
References in corpus (11)
- Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
- Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
- How to Escape Saddle Points Efficiently
- Explorations on high dimensional landscapes
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent Kernel
- Weight-space symmetry in deep networks gives rise to permutation saddles, connected by equal-loss valleys across the loss landscape
- Neural Mechanics: Symmetry and Broken Conservation Laws in Deep Learning Dynamics
- On Connected Sublevel Sets in Deep Learning
- Noether: The More Things Change, the More Stay the Same
- The critical locus of overparameterized neural networks
- GENNI: Visualising the Geometry of Equivalences for Neural Network Identifiability
Cited by in corpus (5)
- Unification of Symmetries Inside Neural Networks: Transformer, Feedforward and Neural ODE
- Neuronal diversity can improve machine learning for physics and beyond
- Some observations on the ambivalent role of symmetries in Bayesian inference problems
- Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural Networks
- Noether's Learning Dynamics: Role of Symmetry Breaking in Neural Networks