Stochastic Optimization with Variance Reduction for Infinite Datasets with Finite-Sum Structure
arXiv:1610.00970
Abstract
Stochastic optimization algorithms with variance reduction have proven successful for minimizing large finite sums of functions. Unfortunately, these techniques are unable to deal with stochastic perturbations of input data, induced for example by data augmentation. In such cases, the objective is no longer a finite sum, and the main candidate for optimization is the stochastic gradient descent method (SGD). In this paper, we introduce a variance reduction approach for these settings when the objective is composite and strongly convex. The convergence rate outperforms SGD with a typically much smaller constant factor, which depends on the variance of gradient estimates only due to perturbations on a single example.
Advances in Neural Information Processing Systems (NIPS), Dec 2017, Long Beach, CA, United States
Cited by in corpus (8)
- Federated Variance-Reduced Stochastic Gradient Descent with Robustness to Byzantine Attacks
- Estimate Sequences for Stochastic Composite Optimization: Variance Reduction, Acceleration, and Robustness to Noise
- Stochastic Newton and Cubic Newton Methods with Simple Local Linear-Quadratic Rates
- A Generic Acceleration Framework for Stochastic Composite Optimization
- On the Ineffectiveness of Variance Reduced Optimization for Deep Learning
- Lower Bounds for Smooth Nonconvex Finite-Sum Optimization
- MISO is Making a Comeback With Better Proofs and Rates
- Variance Reduction in Deep Learning: More Momentum is All You Need