Gradient descent aligns the layers of deep linear networks
arXiv:1810.02032
Abstract
This paper establishes risk convergence and asymptotic weight matrix alignment --- a form of implicit regularization --- of gradient flow and gradient descent when applied to deep linear networks on linearly separable data. In more detail, for gradient flow applied to strictly decreasing loss functions (with similar results for gradient descent with particular decreasing step sizes): (i) the risk converges to 0; (ii) the normalized i-th weight matrix asymptotically equals its rank-1 approximation ; (iii) these rank-1 matrices are aligned across layers, meaning . In the case of the logistic loss (binary cross entropy), more can be said: the linear function induced by the network --- the product of its weight matrices --- converges to the same direction as the maximum margin solution. This last property was identified in prior work, but only under assumptions on gradient descent which here are implied by the alignment phenomenon.
Cited by in corpus (67)
- Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
- Optimization for deep learning: theory and algorithms
- The Modern Mathematics of Deep Learning
- The Convergence Rate of Neural Networks for Learned Functions of Different Frequencies
- Gradient Descent Maximizes the Margin of Homogeneous Neural Networks
- Understanding Dimensional Collapse in Contrastive Self-supervised Learning
- Implicit Regularization in Deep Matrix Factorization
- Implicit Regularization in Deep Learning May Not Be Explainable by Norms
- Self-supervised Learning is More Robust to Dataset Imbalance
- Kernel and Rich Regimes in Overparametrized Models
- Deep Learning Theory Review: An Optimal Control and Dynamical Systems Perspective
- Generalization Guarantees for Neural Networks via Harnessing the Low-rank Structure of the Jacobian
- Width Provably Matters in Optimization for Deep Linear Neural Networks
- Lexicographic and Depth-Sensitive Margins in Homogeneous and Non-Homogeneous Deep Models
- Exponential Convergence Time of Gradient Descent for One-Dimensional Deep Linear Neural Networks
- Shape Matters: Understanding the Implicit Bias of the Noise Covariance
- Revealing the Structure of Deep Neural Networks via Convex Duality
- Directional convergence and alignment in deep learning
- Understanding the role of importance weighting for deep learning
- Towards Resolving the Implicit Bias of Gradient Descent for Matrix Factorization: Greedy Low-Rank Learning
- The Surprising Simplicity of the Early-Time Learning Dynamics of Neural Networks
- A Unifying View on Implicit Bias in Training Linear Neural Networks
- Compression based bound for non-compressed network: unified generalization error analysis of large compressible deep neural network
- On Dropout and Nuclear Norm Regularization
- On the Global Convergence of Training Deep Linear ResNets
- Inductive Bias of Gradient Descent based Adversarial Training on Separable Data
- Understanding Adversarial Robustness: The Trade-off between Minimum and Average Margin
- On Connected Sublevel Sets in Deep Learning
- Implicit Bias in Deep Linear Classification: Initialization Scale vs Training Accuracy
- On the Implicit Bias of Initialization Shape: Beyond Infinitesimal Mirror Descent
- Truth or Backpropaganda? An Empirical Investigation of Deep Learning Theory
- Implicit Regularization in ReLU Networks with the Square Loss
- To Each Optimizer a Norm, To Each Norm its Generalization
- Inductive Bias of Multi-Channel Linear Convolutional Networks with Bounded Weight Norm
- When does gradient descent with logistic loss interpolate using deep networks with smoothed ReLU activations?
- A Theoretical Analysis of Fine-tuning with Linear Teachers
- Global Convergence of Gradient Descent for Deep Linear Residual Networks
- Large Learning Rate Tames Homogeneity: Convergence and Balancing Effect
- Continuous vs. Discrete Optimization of Deep Neural Networks
- When does gradient descent with logistic loss find interpolating two-layer networks?
- Implicit Bias of SGD for Diagonal Linear Networks: a Provable Benefit of Stochasticity
- Global Convergence and Generalization Bound of Gradient-Based Meta-Learning with Deep Neural Nets
- Gradient descent follows the regularization path for general losses
- The Implicit Bias for Adaptive Optimization Algorithms on Homogeneous Neural Networks
- An Optimization and Generalization Analysis for Max-Pooling Networks
- Bridging the Gap Between Adversarial Robustness and Optimization Bias
- Implicit bias of deep linear networks in the large learning rate phase
- Noisy Gradient Descent Converges to Flat Minima for Nonconvex Matrix Factorization
- Implicit Regularization in Tensor Factorization
- On Alignment in Deep Linear Neural Networks
- Properties of the After Kernel
- Interpolation can hurt robust generalization even when there is no noise
- On Margin Maximization in Linear and ReLU Networks
- Unique Properties of Flat Minima in Deep Networks
- Achieving Small Test Error in Mildly Overparameterized Neural Networks
- A Modular Analysis of Provable Acceleration via Polyak's Momentum: Training a Wide ReLU Network and a Deep Linear Network
- Implicitly Maximizing Margins with the Hinge Loss
- Training Linear Neural Networks: Non-Local Convergence and Complexity Results
- Training Efficiency and Robustness in Deep Learning
- Understanding Modern Techniques in Optimization: Frank-Wolfe, Nesterov's Momentum, and Polyak's Momentum
- Deep orthogonal linear networks are shallow
- Directional Convergence Analysis under Spherically Symmetric Distribution
- Gradient Methods Never Overfit On Separable Data
- Recent Advances in Large Margin Learning
- Implicit Regularization of Bregman Proximal Point Algorithm and Mirror Descent on Separable Data
- MSR-DARTS: Minimum Stable Rank of Differentiable Architecture Search
- Towards Understanding Learning in Neural Networks with Linear Teachers