Big Batch SGD: Automated Inference using Adaptive Batch Sizes
arXiv:1610.05792
Abstract
Classical stochastic gradient methods for optimization rely on noisy gradient approximations that become progressively less accurate as iterates approach a solution. The large noise and small signal in the resulting gradients makes it difficult to use them for adaptive stepsize selection and automatic stopping. We propose alternative "big batch" SGD schemes that adaptively grow the batch size over time to maintain a nearly constant signal-to-noise ratio in the gradient approximation. The resulting methods have similar convergence rates to classical SGD, and do not require convexity of the objective. The high fidelity gradients enable automated learning rate selection and do not require stepsize decay. Big batch methods are thus easily automated and can run with little or no oversight.
A preliminary version of this paper appears in AISTATS 2017 (International Conference on Artificial Intelligence and Statistics)
References in corpus (8)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- ADADELTA: An Adaptive Learning Rate Method
- SAGA: A Fast Incremental Gradient Method With Support for Non-Strongly Convex Composite Objectives
- No More Pesky Learning Rates
- Barzilai-Borwein Step Size for Stochastic Gradient Descent
- On Variance Reduction in Stochastic Gradient Descent and its Asynchronous Variants
- Finito: A Faster, Permutable Incremental Gradient Method for Big Data Problems
- Online Learning to Sample
Cited by in corpus (17)
- An Empirical Model of Large-Batch Training
- AdaBatch: Adaptive Batch Sizes for Training Deep Neural Networks
- Advances in Variational Inference
- Painless Stochastic Gradient: Interpolation, Line-Search, and Convergence Rates
- Global Convergence of Arbitrary-Block Gradient Methods for Generalized Polyak-Łojasiewicz Functions
- History-Gradient Aided Batch Size Adaptation for Variance Reduced Algorithms
- Accelerating Minibatch Stochastic Gradient Descent using Typicality Sampling
- Fast and Furious Convergence: Stochastic Second Order Methods under Interpolation
- Linear Range in Gradient Descent
- Improving the convergence of SGD through adaptive batch sizes
- Drill the Cork of Information Bottleneck by Inputting the Most Important Data
- Data Sampling Strategies in Stochastic Algorithms for Empirical Risk Minimization
- Parabolic Approximation Line Search for DNNs
- A Resizable Mini-batch Gradient Descent based on a Multi-Armed Bandit
- Flexible numerical optimization with ensmallen
- Using a one dimensional parabolic model of the full-batch loss to estimate learning rates during training
- Data-driven Algorithm Selection and Parameter Tuning: Two Case studies in Optimization and Signal Processing