From Symmetry to Geometry: Tractable Nonconvex Problems
arXiv:2007.06753
Abstract
As science and engineering have become increasingly data-driven, the role of optimization has expanded to touch almost every stage of the data analysis pipeline, from signal and data acquisition to modeling and prediction. The optimization problems encountered in practice are often nonconvex. While challenges vary from problem to problem, one common source of nonconvexity is nonlinearity in the data or measurement model. Nonlinear models often exhibit symmetries, creating complicated, nonconvex objective landscapes, with multiple equivalent solutions. Nevertheless, simple methods (e.g., gradient descent) often perform surprisingly well in practice. The goal of this survey is to highlight a class of tractable nonconvex problems, which can be understood through the lens of symmetries. These problems exhibit a characteristic geometric structure: local minimizers are symmetric copies of a single "ground truth" solution, while other critical points occur at balanced superpositions of symmetric copies of the ground truth, and exhibit negative curvature in directions that break the symmetry. This structure enables efficient methods to obtain global minimizers. We discuss examples of this phenomenon arising from a wide range of problems in imaging, signal processing, and data analysis. We highlight the key role of symmetry in shaping the objective landscape and discuss the different roles of rotational and discrete symmetries. This area is rich with observed phenomena and open problems; we close by highlighting directions for future research.
review paper, 38 pages, 10 figures, revision: correction of typos, adding more discussion on recent advances on deep learning
References in corpus (19)
- Sparsity and Incoherence in Compressive Sampling
- Robust PCA via Outlier Pursuit
- Optimization for deep learning: theory and algorithms
- Learning One-hidden-layer Neural Networks with Landscape Design
- Exact Recovery of Sparsely-Used Dictionaries
- Mathematics of Deep Learning
- First-order Methods Almost Always Avoid Saddle Points
- Global analysis of Expectation Maximization for mixtures of two Gaussians
- Accelerated Gradient Descent Escapes Saddle Points Faster than Gradient Descent
- Local Maxima in the Likelihood of Gaussian Mixture Models: Structural Results and Algorithmic Consequences
- Solving SDPs for synchronization and MaxCut problems via the Grothendieck inequality
- Porcupine Neural Networks: (Almost) All Local Optima are Global
- Composite optimization for robust blind deconvolution
- Escaping from saddle points on Riemannian manifolds
- Finding the Sparsest Vectors in a Subspace: Theory, Algorithms, and Applications
- When Does Non-Orthogonal Tensor Decomposition Have No Spurious Local Minima?
- Analysis of the Optimization Landscapes for Overcomplete Representation Learning
- Structures of Spurious Local Minima in -means
- A Brief Introduction to Manifold Optimization
Cited by in corpus (12)
- A Geometric Analysis of Neural Collapse with Unconstrained Features
- HePPCAT: Probabilistic PCA for Data with Heteroscedastic Noise
- Scaling and Scalability: Provable Nonconvex Low-Rank Tensor Estimation from Incomplete Measurements
- An Unconstrained Layer-Peeled Perspective on Neural Collapse
- Unique sparse decomposition of low rank matrices
- Convex and Nonconvex Optimization Are Both Minimax-Optimal for Noisy Blind Deconvolution under Random Designs
- The loss landscape of deep linear neural networks: a second-order analysis
- Rank Overspecified Robust Matrix Recovery: Subgradient Method and Exact Recovery
- Nonconvex Factorization and Manifold Formulations are Almost Equivalent in Low-rank Matrix Optimization
- Lecture notes on non-convex algorithms for low-rank matrix recovery
- Sharp global convergence guarantees for iterative nonconvex optimization: A Gaussian process perspective
- Learning Mixtures of Low-Rank Models