Statistical limits of dictionary learning: random matrix theory and the spectral replica method
arXiv:2109.06610 · doi:10.1103/PhysRevE.106.024136
Abstract
We consider increasingly complex models of matrix denoising and dictionary learning in the Bayes-optimal setting, in the challenging regime where the matrices to infer have a rank growing linearly with the system size. This is in contrast with most existing literature concerned with the low-rank (i.e., constant-rank) regime. We first consider a class of rotationally invariant matrix denoising problems whose mutual information and minimum mean-square error are computable using techniques from random matrix theory. Next, we analyze the more challenging models of dictionary learning. To do so we introduce a novel combination of the replica method from statistical mechanics together with random matrix theory, coined spectral replica method. This allows us to derive variational formulas for the mutual information between hidden representations and the noisy data of the dictionary learning problem, as well as for the overlaps quantifying the optimal reconstruction error. The proposed method reduces the number of degrees of freedom from matrix entries to eigenvalues (or singular values), and yields Coulomb gas representations of the mutual information which are reminiscent of matrix models in physics. The main ingredients are a combination of large deviation results for random matrices together with a new replica symmetric decoupling ansatz at the level of the probability distributions of eigenvalues (or singular values) of certain overlap matrices and the use of HarishChandra-Itzykson-Zuber spherical integrals.
References in corpus (13)
- Introduction to Random Matrices - Theory and Practice
- Sparse Principal Components Analysis
- From non-ergodic eigenvectors to local resolvent statistics and back: a random matrix perspective
- The Lévy-Rosenzweig-Porter random matrix ensemble
- On the Asymptotic Spectrum of Products of Independent Random Matrices
- Perturbative construction of mean-field equations in extensive-rank matrix factorization and denoising
- An inference problem in a mismatched setting: a spin-glass model with Mattis interaction
- Finite Size Corrections and Likelihood Ratio Fluctuations in the Spiked Wigner Model
- On the large N limit of matrix integrals over the orthogonal group
- Large Deviations Asymptotics of Rectangular Spherical Integral
- Mismatched Estimation of rank-one symmetric matrices under Gaussian noise
- Rank-one matrix estimation: analytic time evolution of gradient descent dynamics
- Noisy Gradient Descent Converges to Flat Minima for Nonconvex Matrix Factorization
Cited by in corpus (13)
- Perturbative construction of mean-field equations in extensive-rank matrix factorization and denoising
- Depth induces scale-averaging in overparameterized linear Bayesian neural networks
- Matrix factorization with neural networks
- Singular Vectors of Sums of Rectangular Random Matrices and Optimal Estimators of High-Rank Signals: The Extensive Spike Model
- Matrix denoising: Bayes-optimal estimators via low-degree polynomials
- Phase transitions induced by standard and predetermined measurements in transmon arrays
- Bayesian reconstruction of memories stored in neural networks from their connectivity
- Spherical Integrals of Sublinear Rank
- Some observations on the ambivalent role of symmetries in Bayesian inference problems
- Bilinear Sequence Regression: A Model for Learning from Long Sequences of High-dimensional Tokens
- Optimal generalisation and learning transition in extensive-width shallow neural networks near interpolation
- Statistical mechanics of extensive-width Bayesian neural networks near interpolation
- Statistical physics of deep learning: Optimal learning of a multi-layer perceptron near interpolation