Algorithmic stability and hypothesis complexity
arXiv:1702.08712
Abstract
We introduce a notion of algorithmic stability of learning algorithms---that we term \emph{argument stability}---that captures stability of the hypothesis output by the learning algorithm in the normed space of functions from which hypotheses are selected. The main result of the paper bounds the generalization error of any learning algorithm in terms of its argument stability. The bounds are based on martingale inequalities in the Banach space to which the hypotheses belong. We apply the general bounds to bound the performance of some learning algorithms based on empirical risk minimization and stochastic gradient descent.
Cited by in corpus (19)
- Implicit Regularization in Nonconvex Statistical Estimation: Gradient Descent Converges Linearly for Phase Retrieval, Matrix Completion, and Blind Deconvolution
- An Information-Theoretic View for Deep Learning
- Gradient Diversity: a Key Ingredient for Scalable Distributed Learning
- Stability and Generalization of Learning Algorithms that Converge to Global Optima
- Stability of Stochastic Gradient Descent on Nonsmooth Convex Losses
- Speedy Performance Estimation for Neural Architecture Search
- Fine-Grained Analysis of Stability and Generalization for Stochastic Gradient Descent
- Curvature of Feasible Sets in Offline and Online Optimization
- Deep Active Learning by Leveraging Training Dynamics
- Orthogonal Deep Neural Networks
- Stability and Generalization of Stochastic Gradient Methods for Minimax Problems
- Local and Global Uniform Convexity Conditions
- Differentially Private SGD with Non-Smooth Losses
- An Optimal Transport View on Generalization
- Improved Learning Rates for Stochastic Optimization
- On the Rates of Convergence from Surrogate Risk Minimizers to the Bayes Optimal Classifier
- Bayesian Counterfactual Risk Minimization
- Stochastic Gradient Descent with Exponential Convergence Rates of Expected Classification Errors
- Tangent Space Sensitivity and Distribution of Linear Regions in ReLU Networks