Limitations of Implicit Bias in Matrix Sensing: Initialization Rank Matters
arXiv:2008.12091
Abstract
In matrix sensing, we first numerically identify the sensitivity to the initialization rank as a new limitation of the implicit bias of gradient flow. We will partially quantify this phenomenon mathematically, where we establish that the gradient flow of the empirical risk is implicitly biased towards low-rank outcomes and successfully learns the planted low-rank matrix, provided that the initialization is low-rank and within a specific "capture neighborhood". This capture neighborhood is far larger than the corresponding neighborhood in local refinement results; the former contains all models with zero training error whereas the latter is a small neighborhood of a model with zero test error. These new insights enable us to design an alternative algorithm for matrix sensing that complements the high-rank and near-zero initialization scheme which is predominant in the existing literature.
References in corpus (5)
- Nonconvex Optimization Meets Low-Rank Matrix Factorization: An Overview
- Implicit Regularization in Deep Learning May Not Be Explainable by Norms
- Implicit Regularization of Discrete Gradient Dynamics in Linear Neural Networks
- Can Implicit Bias Explain Generalization? Stochastic Convex Optimization as a Case Study
- Implicit Regularization in ReLU Networks with the Square Loss