Lexicographic and Depth-Sensitive Margins in Homogeneous and Non-Homogeneous Deep Models
arXiv:1905.07325
Abstract
With an eye toward understanding complexity control in deep learning, we study how infinitesimal regularization or gradient descent optimization lead to margin maximizing solutions in both homogeneous and non-homogeneous models, extending previous work that focused on infinitesimal regularization only in homogeneous models. To this end we study the limit of loss minimization with a diverging norm constraint (the "constrained path"), relate it to the limit of a "margin path" and characterize the resulting solution. For non-homogeneous ensemble models, which output is a sum of homogeneous sub-models, we show that this solution discards the shallowest sub-models if they are unnecessary. For homogeneous models, we show convergence to a "lexicographic max-margin solution", and provide conditions under which max-margin solutions are also attained as the limit of unconstrained gradient descent.
ICML Camera ready version
References in corpus (1)
Cited by in corpus (9)
- Shape Matters: Understanding the Implicit Bias of the Noise Covariance
- A Function Space View of Bounded Norm Infinite Width ReLU Nets: The Multivariate Case
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural Networks
- Understanding the role of importance weighting for deep learning
- Explicit regularization and implicit bias in deep network classifiers trained with the square loss
- Inductive Bias of Multi-Channel Linear Convolutional Networks with Bounded Weight Norm
- Implicit bias of deep linear networks in the large learning rate phase
- Distribution of Classification Margins: Are All Data Equal?
- Going Beyond Linear RL: Sample Efficient Neural Function Approximation