Implicit Bias of Gradient Descent on Linear Convolutional Networks
arXiv:1806.00468
Abstract
We show that gradient descent on full-width linear convolutional networks of depth converges to a linear predictor related to the bridge penalty in the frequency domain. This is in contrast to linearly fully connected networks, where gradient descent converges to the hard margin linear support vector machine solution, regardless of depth.
References in corpus (11)
- Don't Decay the Learning Rate, Increase the Batch Size
- Low-rank optimization for semidefinite convex problems
- The loss surface of deep and wide neural networks
- Characterizing Implicit Bias in Terms of Optimization Geometry
- Path-SGD: Path-Normalized Optimization in Deep Neural Networks
- Sharp Minima Can Generalize For Deep Nets
- Entropy-SGD: Biasing Gradient Descent Into Wide Valleys
- Geometry of Optimization and Implicit Regularization in Deep Learning
- Risk and parameter convergence of logistic regression
- Margins, Shrinkage, and Boosting
- Convergence of Gradient Descent on Separable Data
Cited by in corpus (16)
- Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
- A Generalized Neural Tangent Kernel Analysis for Two-layer Neural Networks
- How Much Over-parameterization Is Sufficient to Learn Deep ReLU Networks?
- Convergence of Gradient Descent on Separable Data
- Lexicographic and Depth-Sensitive Margins in Homogeneous and Non-Homogeneous Deep Models
- Implicit Regularization of Discrete Gradient Dynamics in Linear Neural Networks
- Revealing the Structure of Deep Neural Networks via Convex Duality
- Small random initialization is akin to spectral learning: Optimization and generalization guarantees for overparameterized low-rank matrix reconstruction
- On Dropout and Nuclear Norm Regularization
- Analytic Network Learning
- Towards Understanding the Generalization Bias of Two Layer Convolutional Linear Classifiers with Gradient Descent
- An Optimization and Generalization Analysis for Max-Pooling Networks
- Understanding Deflation Process in Over-parametrized Tensor Decomposition
- An Upper Limit of Decaying Rate with Respect to Frequency in Deep Neural Network
- Interpolation can hurt robust generalization even when there is no noise
- Logit Attenuating Weight Normalization