Variance Reduction in SGD by Distributed Importance Sampling
arXiv:1511.06481
Abstract
Humans are able to accelerate their learning by selecting training materials that are the most informative and at the appropriate level of difficulty. We propose a framework for distributing deep learning in which one set of workers search for the most informative examples in parallel while a single worker updates the model on examples selected by importance sampling. This leads the model to update using an unbiased estimate of the gradient which also has minimum variance when the sampling proposal is proportional to the L2-norm of the gradient. We show experimentally that this method reduces gradient variance even in a context where the cost of synchronization across machines cannot be ignored, and where the factors for importance sampling are not updated instantly across the training set.
References in corpus (1)
Cited by in corpus (43)
- Distributed Prioritized Experience Replay
- Efficient training of physics-informed neural networks via importance sampling
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval
- Online Batch Selection for Faster Training of Neural Networks
- Active Bias: Training More Accurate Neural Networks by Emphasizing High Variance Samples
- Not All Samples Are Created Equal: Deep Learning with Importance Sampling
- Biased Importance Sampling for Deep Neural Network Training
- FairBatch: Batch Selection for Model Fairness
- Accelerating Deep Learning by Focusing on the Biggest Losers
- An Equivalence between Loss Functions and Non-Uniform Sampling in Experience Replay
- Variance-reduced Language Pretraining via a Mask Proposal Network
- Adaptive Sketch-and-Project Methods for Solving Linear Systems
- A Survey on Curriculum Learning
- Federated Transfer Learning with Dynamic Gradient Aggregation
- Derivative Manipulation for General Example Weighting
- Active Mini-Batch Sampling using Repulsive Point Processes
- PANDA: Facilitating Usable AI Development
- Accelerating Minibatch Stochastic Gradient Descent using Typicality Sampling
- Efficient Per-Example Gradient Computations in Convolutional Neural Networks
- AutoAssist: A Framework to Accelerate Training of Deep Neural Networks
- Online Variance Reduction for Stochastic Optimization
- Adaptive Task Sampling for Meta-Learning
- Adaptive Sampling Distributed Stochastic Variance Reduced Gradient for Heterogeneous Distributed Datasets
- Gradient Descent in RKHS with Importance Labeling
- Lsh-sampling Breaks the Computation Chicken-and-egg Loop in Adaptive Stochastic Gradient Estimation
- Stochastic Reweighted Gradient Descent
- Disentangling Sampling and Labeling Bias for Learning in Large-Output Spaces
- Modern Subsampling Methods for Large-Scale Least Squares Regression
- Improving Self-supervised Pre-training via a Fully-Explored Masked Language Model
- A Survey on Large-scale Machine Learning
- Selective sampling for accelerating training of deep neural networks
- Margin-Based Regularization and Selective Sampling in Deep Neural Networks
- Drill the Cork of Information Bottleneck by Inputting the Most Important Data
- Adam with Bandit Sampling for Deep Learning
- Learning to Auto Weight: Entirely Data-driven and Highly Efficient Weighting Framework
- Label and Sample: Efficient Training of Vehicle Object Detector from Sparsely Labeled Data
- Using Wavelets to Analyze Similarities in Image-Classification Datasets
- One Backward from Ten Forward, Subsampling for Large-Scale Deep Learning
- Submodular Batch Selection for Training Deep Neural Networks
- Dynamic Gradient Aggregation for Federated Domain Adaptation
- Efficient Reinforcement Learning in Resource Allocation Problems Through Permutation Invariant Multi-task Learning
- Accelerate RNN-based Training with Importance Sampling
- Training Efficiency and Robustness in Deep Learning