On the Convergence of Nested Decentralized Gradient Methods with Multiple Consensus and Gradient Steps
arXiv:2006.01665 · doi:10.1109/TSP.2021.3094906
Abstract
In this paper, we consider minimizing a sum of local convex objective functions in a distributed setting, where the cost of communication and/or computation can be expensive. We extend and generalize the analysis for a class of nested gradient-based distributed algorithms (NEAR-DGD; Berahas, Bollapragada, Keskar and Wei, 2018) to account for multiple gradient steps at every iteration. We show the effect of performing multiple gradient steps on the rate of convergence and on the size of the neighborhood of convergence, and prove R-Linear convergence to the exact solution with a fixed number of gradient steps and increasing number of consensus steps. We test the performance of the generalized method on quadratic functions and show the effect of multiple consensus and gradient steps in terms of iterations, number of gradient evaluations, number of communications and cost.
12 pages, 4 figures. arXiv admin note: text overlap with arXiv:1903.08149
References in corpus (7)
- Federated Learning: Challenges, Methods, and Future Directions
- Towards Federated Learning at Scale: System Design
- Decentralized Stochastic Optimization and Gossip Algorithms with Compressed Communication
- FedSplit: An algorithmic framework for fast federated optimization
- Parallel SGD: When does averaging help?
- Communication/Computation Tradeoffs in Consensus-Based Distributed Optimization
- Multi-consensus Decentralized Accelerated Gradient Descent