Asynchronous Parallel Stochastic Gradient for Nonconvex Optimization
arXiv:1506.08272
Abstract
Asynchronous parallel implementations of stochastic gradient (SG) have been broadly used in solving deep neural network and received many successes in practice recently. However, existing theories cannot explain their convergence and speedup properties, mainly due to the nonconvexity of most deep learning formulations and the asynchronous parallel mechanism. To fill the gaps in theory and provide theoretical supports, this paper studies two asynchronous parallel implementations of SG: one is on the computer network and the other is on the shared memory system. We establish an ergodic convergence rate for both algorithms and prove that the linear speedup is achievable if the number of workers is bounded by ( is the total number of iterations). Our results generalize and improve existing analysis for convex minimization.
33 pages
References in corpus (1)
Cited by in corpus (101)
- Can Decentralized Algorithms Outperform Centralized Algorithms? A Case Study for Decentralized Parallel Stochastic Gradient Descent
- Stochastic Variance Reduction for Nonconvex Optimization
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- Local SGD Converges Fast and Communicates Little
- MD-GAN: Multi-Discriminator Generative Adversarial Networks for Distributed Datasets
- Gradient Sparsification for Communication-Efficient Distributed Optimization
- Staleness-aware Async-SGD for Distributed Deep Learning
- Cooperative SGD: A unified Framework for the Design and Analysis of Communication-Efficient SGD Algorithms
- Asynchronous Stochastic Gradient Descent with Delay Compensation
- AdaNet: Adaptive Structural Learning of Artificial Neural Networks
- Communication-Efficient Distributed Deep Learning: A Comprehensive Survey
- Variance Reduction in SGD by Distributed Importance Sampling
- Adaptive Federated Learning in Resource Constrained Edge Computing Systems
- Towards Efficient and Stable K-Asynchronous Federated Learning with Unbounded Stale Gradients on Non-IID Data
- Efficient and Robust Parallel DNN Training through Model Parallelism on Multi-GPU Platform
- The Error-Feedback Framework: Better Rates for SGD with Delayed Gradients and Compressed Communication
- Asynchronous Decentralized Parallel Stochastic Gradient Descent
- HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed Framework
- Omnivore: An Optimizer for Multi-device Deep Learning on CPUs and GPUs
- Asynchronous and Parallel Distributed Pose Graph Optimization
- DataLens: Scalable Privacy Preserving Training via Gradient Compression and Aggregation
- Improved asynchronous parallel optimization analysis for stochastic incremental methods
- Semi-Synchronous Federated Learning for Energy-Efficient Training and Accelerated Convergence in Cross-Silo Settings
- Federated Learning with Buffered Asynchronous Aggregation
- CYCLADES: Conflict-free Asynchronous Machine Learning
- A Double Residual Compression Algorithm for Efficient Distributed Learning
- The Sound of APALM Clapping: Faster Nonsmooth Nonconvex Optimization with Stochastic Asynchronous PALM
- Bayesian Pose Graph Optimization via Bingham Distributions and Tempered Geodesic MCMC
- Semi-Synchronous Personalized Federated Learning over Mobile Edge Networks
- Decentralized Stochastic Gradient Tracking for Non-convex Empirical Risk Minimization
- The Asynchronous PALM Algorithm for Nonsmooth Nonconvex Problems
- Training Large Neural Networks with Constant Memory using a New Execution Algorithm
- Byzantine-Tolerant Machine Learning
- Asynchronous Stochastic Gradient Descent with Variance Reduction for Non-Convex Optimization
- MXNET-MPI: Embedding MPI parallelism in Parameter Server Task Model for scaling Deep Learning
- Communication trade-offs for synchronized distributed SGD with large step size
- Achieving Linear Convergence in Distributed Asynchronous Multi-agent Optimization
- Faster Asynchronous SGD
- Distributed Momentum for Byzantine-resilient Learning
- DeepSpark: A Spark-Based Distributed Deep Learning Framework for Commodity Clusters
- The Convergence of Sparsified Gradient Methods
- Handover Control in Wireless Systems via Asynchronous Multi-User Deep Reinforcement Learning
- Asynchronous Parallel Algorithms for Nonconvex Big-Data Optimization. Part II: Complexity and Numerical Results
- DataBright: Towards a Global Exchange for Decentralized Data Ownership and Trusted Computation
- An introduction to decentralized stochastic optimization with gradient tracking
- SparCML: High-Performance Sparse Communication for Machine Learning
- Fully Decoupled Neural Network Learning Using Delayed Gradients
- On Nonconvex Decentralized Gradient Descent
- Zeroth-order Asynchronous Doubly Stochastic Algorithm with Variance Reduction
- Secure Bilevel Asynchronous Vertical Federated Learning with Backward Updating
- Pipelined Backpropagation at Scale: Training Large Models without Batches
- Distributed Deep Learning with Event-Triggered Communication
- EventGraD: Event-Triggered Communication in Parallel Machine Learning
- Asynchronous Stochastic Block Coordinate Descent with Variance Reduction
- Parareal Neural Networks Emulating a Parallel-in-time Algorithm
- Accelerating Asynchronous Algorithms for Convex Optimization by Momentum Compensation
- At Stability's Edge: How to Adjust Hyperparameters to Preserve Minima Selection in Asynchronous Training of Neural Networks?
- Elastic Consistency: A General Consistency Model for Distributed Stochastic Gradient Descent
- DBS: Dynamic Batch Size For Distributed Deep Neural Network Training
- Async-RED: A Provably Convergent Asynchronous Block Parallel Stochastic Method using Deep Denoising Priors
- Taming Convergence for Asynchronous Stochastic Gradient Descent with Unbounded Delay in Non-Convex Learning
- Communication-Efficient Federated Learning with Compensated Overlap-FedAvg
- Private and Communication-Efficient Edge Learning: A Sparse Differential Gaussian-Masking Distributed SGD Approach
- NUQSGD: Provably Communication-efficient Data-parallel SGD via Nonuniform Quantization
- Distributed Asynchronous Dual Free Stochastic Dual Coordinate Ascent
- Asynchronous Stochastic Proximal Methods for Nonconvex Nonsmooth Optimization
- Faster Distributed Deep Net Training: Computation and Communication Decoupled Stochastic Gradient Descent
- Make Workers Work Harder: Decoupled Asynchronous Proximal Stochastic Gradient Descent
- Improved Learning Rates for Stochastic Optimization
- Stochastic Distributed Optimization for Machine Learning from Decentralized Features
- Asynchronous Parallel Empirical Variance Guided Algorithms for the Thresholding Bandit Problem
- Decoupled Asynchronous Proximal Stochastic Gradient Descent with Variance Reduction
- Citadel: Protecting Data Privacy and Model Confidentiality for Collaborative Learning with SGX
- A flexible framework for communication-efficient machine learning: from HPC to IoT
- Efficient Learning of Generative Models via Finite-Difference Score Matching
- The Convergence of Stochastic Gradient Descent in Asynchronous Shared Memory
- Parallel and distributed asynchronous adaptive stochastic gradient methods
- Step-Ahead Error Feedback for Distributed Training with Compressed Gradient
- Hogwild! over Distributed Local Data Sets with Linearly Increasing Mini-Batch Sizes
- Differential Equations for Modeling Asynchronous Algorithms
- Accelerated Sparsified SGD with Error Feedback
- Asynchronous Optimization Methods for Efficient Training of Deep Neural Networks with Guarantees
- Making Asynchronous Stochastic Gradient Descent Work for Transformers
- Sparsification as a Remedy for Staleness in Distributed Asynchronous SGD
- Asynchronous Iterations in Optimization: New Sequence Results and Sharper Algorithmic Guarantees
- Good Intentions: Adaptive Parameter Management via Intent Signaling
- New Aspects of Black Box Conditional Gradient: Variance Reduction and One Point Feedback
- Bounding the expected run-time of nonconvex optimization with early stopping
- Distributed deep learning on edge-devices: feasibility via adaptive compression
- Gap Aware Mitigation of Gradient Staleness
- Asynchronous Stochastic Optimization Robust to Arbitrary Delays
- Accumulated Decoupled Learning: Mitigating Gradient Staleness in Inter-Layer Model Parallelization
- Fully Distributed and Asynchronized Stochastic Gradient Descent for Networked Systems
- The Gradient Convergence Bound of Federated Multi-Agent Reinforcement Learning with Efficient Communication
- Tell Me Something New: A New Framework for Asynchronous Parallel Learning
- GaDei: On Scale-up Training As A Service For Deep Learning
- Towards Understanding Acceleration Tradeoff between Momentum and Asynchrony in Nonconvex Stochastic Optimization
- One-Point Feedback for Composite Optimization with Applications to Distributed and Federated Learning
- Distributed Networked Learning with Correlated Data
- A Model Parallel Proximal Stochastic Gradient Algorithm for Partially Asynchronous Systems
- Asynchronous Parallel Nonconvex Optimization Under the Polyak-Lojasiewicz Condition