Improving the Sample and Communication Complexity for Decentralized Non-Convex Optimization: A Joint Gradient Estimation and Tracking Approach
arXiv:1910.05857
Abstract
Many modern large-scale machine learning problems benefit from decentralized and stochastic optimization. Recent works have shown that utilizing both decentralized computing and local stochastic gradient estimates can outperform state-of-the-art centralized algorithms, in applications involving highly non-convex problems, such as training deep neural networks. In this work, we propose a decentralized stochastic algorithm to deal with certain smooth non-convex problems where there are nodes in the system, and each node has a large number of samples (denoted as ). Differently from the majority of the existing decentralized learning algorithms for either stochastic or finite-sum problems, our focus is given to both reducing the total communication rounds among the nodes, while accessing the minimum number of local data samples. In particular, we propose an algorithm named D-GET (decentralized gradient estimation and tracking), which jointly performs decentralized gradient estimation (which estimates the local gradient using a subset of local samples) and gradient tracking (which tracks the global full gradient using local estimates). We show that, to achieve certain stationary solution of the deterministic finite sum problem, the proposed algorithm achieves an sample complexity and an communication complexity. These bounds significantly improve upon the best existing bounds of and , respectively. Similarly, for online problems, the proposed method achieves an sample complexity and an communication complexity, while the best existing bounds are and , respectively.
References in corpus (18)
- Federated Learning: Strategies for Improving Communication Efficiency
- On the Linear Convergence of the ADMM in Decentralized Consensus Optimization
- Diffusion Adaptation Strategies for Distributed Optimization and Learning over Networks
- D-ADMM: A Communication-Efficient Distributed Algorithm For Separable Optimization
- SPIDER: Near-Optimal Non-Convex Optimization via Stochastic Path Integrated Differential Estimator
- Optimal algorithms for smooth and strongly convex distributed optimization in networks
- D: Decentralized Training over Decentralized Data
- DSA: Decentralized Double Stochastic Averaging Gradient Algorithm
- Stochastic Gradient Push for Distributed Deep Learning
- Non-convex Finite-Sum Optimization Via SCSG Methods
- Dictionary Learning over Distributed Models
- Optimal Algorithms for Non-Smooth Distributed Optimization in Networks
- Distributed Non-Convex First-Order Optimization and Information Processing: Lower Complexity Bounds and Rate Optimal Algorithms
- Stochastic Learning under Random Reshuffling with Constant Step-sizes
- An Accelerated Decentralized Stochastic Proximal Algorithm for Finite Sums
- Communication-Efficient Distributed Optimization in Networks with Gradient Tracking and Variance Reduction
- Variance-Reduced Decentralized Stochastic Optimization with Gradient Tracking--Part I: GT-SAGA
- Towards More Efficient Stochastic Decentralized Learning: Faster Convergence and Sparse Communication
Cited by in corpus (5)
- Parallel Restarted SPIDER -- Communication Efficient Distributed Nonconvex Optimization with Optimal Computation Complexity
- Communication-Efficient Distributed Optimization in Networks with Gradient Tracking and Variance Reduction
- A fast randomized incremental gradient method for decentralized non-convex optimization
- A general framework for decentralized optimization with first-order methods
- Decentralized Learning with Lazy and Approximate Dual Gradients