Distributed Gradient Methods with Variable Number of Working Nodes
arXiv:1504.04049 · doi:10.1109/TSP.2016.2560133
Abstract
We consider distributed optimization where nodes in a connected network minimize the sum of their local costs subject to a common constraint set. We propose a distributed projected gradient method where each node, at each iteration , performs an update (is active) with probability , and stays idle (is inactive) with probability . Whenever active, each node performs an update by weight-averaging its solution estimate with the estimates of its active neighbors, taking a negative gradient step with respect to its local cost, and performing a projection onto the constraint set; inactive nodes perform no updates. Assuming that nodes' local costs are strongly convex, with Lipschitz continuous gradients, we show that, as long as activation probability grows to one asymptotically, our algorithm converges in the mean square sense (MSS) to the same solution as the standard distributed gradient method, i.e., as if all the nodes were active at all iterations. Moreover, when grows to one linearly, with an appropriately set convergence factor, the algorithm has a linear MSS convergence, with practically the same factor as the standard distributed gradient method. Simulations on both synthetic and real world data sets demonstrate that, when compared with the standard distributed gradient method, the proposed algorithm significantly reduces the overall number of per-node communications and per-node gradient evaluations (computational cost) for the same required accuracy.
submitted to a journal on April 15, 2015; revised on September 23, 2015, and March 10, 2016
References in corpus (6)
- Sensor Networks with Random Links: Topology Design for Distributed Consensus
- Hybrid Random/Deterministic Parallel Algorithms for Nonconvex Big Data Optimization
- Network Newton-Part II: Convergence Rate and Implementation
- Network Newton-Part I: Algorithm and Convergence
- Distributed Gradient Methods with Variable Number of Working Nodes
- Asynchronous Adaptation and Learning over Networks - Part II: Performance Analysis
Cited by in corpus (7)
- Distributed Optimization for Smart Cyber-Physical Networks
- Distributed Gradient Methods with Variable Number of Working Nodes
- Accelerated Distributed Dual Averaging over Evolving Networks of Growing Connectivity
- Communication-Efficient Distributed Strongly Convex Stochastic Optimization: Non-Asymptotic Rates
- Resource-aware Exact Decentralized Optimization Using Event-triggered Broadcasting
- A Unification and Generalization of Exact Distributed First Order Methods
- On the Convergence of Inexact Gradient Descent with Controlled Synchronization Steps