Risk and parameter convergence of logistic regression
arXiv:1803.07300
Abstract
Gradient descent, when applied to the task of logistic regression, outputs iterates which are biased to follow a unique ray defined by the data. The direction of this ray is the maximum margin predictor of a maximal linearly separable subset of the data; the gradient descent iterates converge to this ray in direction at the rate . The ray does not pass through the origin in general, and its offset is the bounded global optimum of the risk over the remaining data; gradient descent recovers this offset at a rate .
Appears in COLT 2019 with the title "The implicit bias of gradient descent on nonseparable data" (and no other changes)
References in corpus (3)
Cited by in corpus (46)
- The Convergence Rate of Neural Networks for Learned Functions of Different Frequencies
- Gradient Descent Maximizes the Margin of Homogeneous Neural Networks
- Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networks
- Generalization Error Bounds of Gradient Descent for Learning Over-parameterized Deep ReLU Networks
- Understanding the Failure Modes of Out-of-Distribution Generalization
- Spurious Valleys in Two-layer Neural Network Optimization Landscapes
- A Selective Overview of Deep Learning
- Generalization Guarantees for Neural Networks via Harnessing the Low-rank Structure of the Jacobian
- Asymmetric Valleys: Beyond Sharp and Flat Local Minima
- Implicit Bias of Gradient Descent on Linear Convolutional Networks
- Improved Sample Complexities for Deep Networks and Robust Classification via an All-Layer Margin
- How Much Over-parameterization Is Sufficient to Learn Deep ReLU Networks?
- Statistical Query Algorithms and Low-Degree Tests Are Almost Equivalent
- Convergence of Gradient Descent on Separable Data
- Lexicographic and Depth-Sensitive Margins in Homogeneous and Non-Homogeneous Deep Models
- Shape Matters: Understanding the Implicit Bias of the Noise Covariance
- Directional convergence and alignment in deep learning
- Understanding the role of importance weighting for deep learning
- Persistency of Excitation for Robustness of Neural Networks
- When Will Gradient Methods Converge to Max-margin Classifier under ReLU Models?
- Convergence and Margin of Adversarial Training on Separable Data
- Label-Imbalanced and Group-Sensitive Classification under Overparameterization
- Large-time asymptotics in deep learning
- Early-stopped neural networks are consistent
- Inductive Bias of Gradient Descent based Adversarial Training on Separable Data
- Robust Large-Margin Learning in Hyperbolic Space
- Towards Understanding the Data Dependency of Mixup-style Training
- Inductive Bias of Multi-Channel Linear Convolutional Networks with Bounded Weight Norm
- Implicit Regularization in ReLU Networks with the Square Loss
- On generalization bounds for deep networks based on loss surface implicit regularization
- Exploring Weight Importance and Hessian Bias in Model Pruning
- Condition Number Analysis of Logistic Regression, and its Implications for Standard First-Order Solution Methods
- An Optimization and Generalization Analysis for Max-Pooling Networks
- Gradient descent follows the regularization path for general losses
- The Dynamics of Gradient Descent for Overparametrized Neural Networks
- Statistical Inference for Polyak-Ruppert Averaged Zeroth-order Stochastic Gradient Algorithm
- A Convergence Theory Towards Practical Over-parameterized Deep Neural Networks
- Actor-critic is implicitly biased towards high entropy optimal policies
- On Margin Maximization in Linear and ReLU Networks
- Understanding Deflation Process in Over-parametrized Tensor Decomposition
- Implicitly Maximizing Margins with the Hinge Loss
- A termination criterion for stochastic gradient descent for binary classification
- Gradient Methods Never Overfit On Separable Data
- Directional Convergence Analysis under Spherically Symmetric Distribution
- Robustifying Binary Classification to Adversarial Perturbation
- SGD: The Role of Implicit Regularization, Batch-size and Multiple-epochs