Understanding Deflation Process in Over-parametrized Tensor Decomposition
arXiv:2106.06573
Abstract
In this paper we study the training dynamics for gradient flow on over-parametrized tensor decomposition problems. Empirically, such training process often first fits larger components and then discovers smaller components, which is similar to a tensor deflation process that is commonly used in tensor decomposition algorithms. We prove that for orthogonally decomposable tensor, a slightly modified version of gradient flow would follow a tensor deflation process and recover all the tensor components. Our proof suggests that for orthogonal tensors, gradient flow dynamics works similarly as greedy low-rank learning in the matrix setting, which is a first step towards understanding the implicit regularization effect of over-parametrized models for low-rank tensors.
NeurIPS 2021 Camera Ready
References in corpus (8)
- Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
- Learning One-hidden-layer Neural Networks with Landscape Design
- Asymptotics of Wide Networks from Feynman Diagrams
- Stochastic Particle Gradient Descent for Infinite Ensembles
- Towards Resolving the Implicit Bias of Gradient Descent for Matrix Factorization: Greedy Low-Rank Learning
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural Networks
- Taylorized Training: Towards Better Approximation of Neural Network Training at Finite Width
- Implicit Bias in Deep Linear Classification: Initialization Scale vs Training Accuracy