On the Principle of Least Symmetry Breaking in Shallow ReLU Models
arXiv:1912.11939
Abstract
We consider the optimization problem associated with fitting two-layer ReLU networks with respect to the squared loss, where labels are assumed to be generated by a target network. Focusing first on standard Gaussian inputs, we show that the structure of spurious local minima detected by stochastic gradient descent (SGD) is, in a well-defined sense, the \emph{least loss of symmetry} with respect to the target weights. A closer look at the analysis indicates that this principle of least symmetry breaking may apply to a broader range of settings. Motivated by this, we conduct a series of experiments which corroborate this hypothesis for different classes of non-isotropic non-product distributions, smooth activation functions and networks with a few layers.
References in corpus (5)
- Recovery Guarantees for One-hidden-layer Neural Networks
- Globally Optimal Gradient Descent for a ConvNet with Gaussian Inputs
- Diverse Neural Network Learns True Target Functions
- Porcupine Neural Networks: (Almost) All Local Optima are Global
- Weight-space symmetry in deep networks gives rise to permutation saddles, connected by equal-loss valleys across the loss landscape
Cited by in corpus (4)
- Symmetry Breaking in Symmetric Tensor Decomposition
- Analytic Characterization of the Hessian in Shallow ReLU Models: A Tale of Symmetry
- The Effects of Mild Over-parameterization on the Optimization Landscape of Shallow ReLU Neural Networks
- Analytic Study of Families of Spurious Minima in Two-Layer ReLU Neural Networks: A Tale of Symmetry II