Reducing Runtime by Recycling Samples
arXiv:1602.02136
Abstract
Contrary to the situation with stochastic gradient descent, we argue that when using stochastic methods with variance reduction, such as SDCA, SAG or SVRG, as well as their variants, it could be beneficial to reuse previously used samples instead of fresh samples, even when fresh samples are available. We demonstrate this empirically for SDCA, SAG and SVRG, studying the optimal sample size one should use, and also uncover be-havior that suggests running SDCA for an integer number of epochs could be wasteful.
References in corpus (5)
- A Stochastic Gradient Method with an Exponential Convergence Rate for Finite Training Sets
- Finito: A Faster, Permutable Incremental Gradient Method for Big Data Problems
- Randomized Dual Coordinate Ascent with Arbitrary Sampling
- Beneath the valley of the noncommutative arithmetic-geometric mean inequality: conjectures, case-studies, and consequences
- Stochastic Dual Coordinate Ascent with Adaptive Probabilities