A Stochastic Operator Framework for Optimization and Learning with Sub-Weibull Errors
arXiv:2105.09884 · doi:10.1109/TAC.2024.3419186
Abstract
This paper proposes a framework to study the convergence of stochastic optimization and learning algorithms. The framework is modeled over the different challenges that these algorithms pose, such as (i) the presence of random additive errors (e.g. due to stochastic gradients), and (ii) random coordinate updates (e.g. due to asynchrony in distributed set-ups). The paper covers both convex and strongly convex problems, and it also analyzes online scenarios, involving changes in the data and costs. The paper relies on interpreting stochastic algorithms as the iterated application of stochastic operators, thus allowing us to use the powerful tools of operator theory. In particular, we consider operators characterized by additive errors with sub-Weibull distribution (which parameterize a broad class of errors by their tail probability), and random updates. In this framework we derive convergence results in mean and in high probability, by providing bounds to the distance of the current iteration from a solution of the optimization or learning problem. The contributions are discussed in light of federated learning applications.
To appear in IEEE Transactions on Automatic Control
References in corpus (21)
- Scikit-learn: Machine Learning in Python
- Federated Learning: Challenges, Methods, and Future Directions
- SCAFFOLD: Stochastic Controlled Averaging for Federated Learning
- ARock: an Algorithmic Framework for Asynchronous Parallel Coordinate Updates
- Federated Learning: A Signal Processing Perspective
- An Exact Quantized Decentralized Gradient Descent Algorithm
- Online Learning with Inexact Proximal Online Gradient Descent Algorithms
- Timescale Separation in Autonomous Optimization
- Online Optimization : Competing with Dynamic Comparators
- Asynchronous Distributed Optimization over Lossy Networks via Relaxed ADMM: Stability and Linear Convergence
- Optimization and Learning with Information Streams: Time-varying Algorithms and Applications
- Coordinate Friendly Structures, Algorithms and Applications
- Moving Beyond Sub-Gaussianity in High-Dimensional Statistics: Applications in Covariance Estimation and Linear Regression
- Concentration Inequalities for Statistical Inference
- The Heavy-Tail Phenomenon in SGD
- Sub-Weibull distributions: generalizing sub-Gaussian and sub-Exponential properties to heavier-tailed distributions
- Understanding Priors in Bayesian Neural Networks at the Unit Level
- SuperMann: a superlinearly convergent algorithm for finding fixed points of nonexpansive operators
- Personalized Optimization with User's Feedback
- tvopt: A Python Framework for Time-Varying Optimization
- High Probability Convergence of Clipped-SGD Under Heavy-tailed Noise