Complete Dictionary Learning via -Norm Maximization over the Orthogonal Group
arXiv:1906.02435
Abstract
This paper considers the fundamental problem of learning a complete (orthogonal) dictionary from samples of sparsely generated signals. Most existing methods solve the dictionary (and sparse representations) based on heuristic algorithms, usually without theoretical guarantees for either optimality or complexity. The recent -minimization based methods do provide such guarantees but the associated algorithms recover the dictionary one column at a time. In this work, we propose a new formulation that maximizes the -norm over the orthogonal group, to learn the entire dictionary. We prove that under a random data model, with nearly minimum sample complexity, the global optima of the norm are very close to signed permutations of the ground truth. Inspired by this observation, we give a conceptually simple and yet effective algorithm based on "matching, stretching, and projection" (MSP). The algorithm provably converges locally at a superlinear (cubic) rate and cost per iteration is merely an SVD. In addition to strong theoretical guarantees, experiments show that the new algorithm is significantly more efficient and effective than existing methods, including KSVD and -based methods. Preliminary experimental results on mixed real imagery data clearly demonstrate advantages of so learned dictionary over classic PCA bases.
References in corpus (12)
- Supervised Dictionary Learning
- A Convergence Theory for Deep Learning via Over-Parameterization
- Generalized power method for sparse principal component analysis
- Nonconvex Optimization Meets Low-Rank Matrix Factorization: An Overview
- Learning Overparameterized Neural Networks via Stochastic Gradient Descent on Structured Data
- Implicit Regularization in Nonconvex Statistical Estimation: Gradient Descent Converges Linearly for Phase Retrieval, Matrix Completion, and Blind Deconvolution
- Exact Recovery of Sparsely-Used Dictionaries
- Sparsifying Transform Learning with Efficient Optimal Updates and Convergence Guarantees
- Simple, Efficient, and Neural Algorithms for Sparse Coding
- Subgradient Descent Learns Orthogonal Dictionaries
- A Nonconvex Approach for Exact and Efficient Multichannel Sparse Blind Deconvolution
- Blind Data Detection in Massive MIMO via -norm Maximization over the Stiefel Manifold
Cited by in corpus (8)
- Nonconvex Optimization Meets Low-Rank Matrix Factorization: An Overview
- Decentralized Optimization Over the Stiefel Manifold by an Approximate Augmented Lagrangian Function
- Weakly Convex Optimization over Stiefel Manifold Using Riemannian Subgradient-Type Methods
- Blind Data Detection in Massive MIMO via -norm Maximization over the Stiefel Manifold
- Complete Dictionary Learning via -norm Maximization
- Finding the Sparsest Vectors in a Subspace: Theory, Algorithms, and Applications
- Manifold Gradient Descent Solves Multi-Channel Sparse Blind Deconvolution Provably and Efficiently
- Brain Image Synthesis with Unsupervised Multivariate Canonical CSCNet