Convergence Rates of Distributed Nesterov-like Gradient Methods on Random Networks
arXiv:1308.0916 · doi:10.1109/TSP.2013.2291221
Abstract
We consider distributed optimization in random networks where N nodes cooperatively minimize the sum \sum_{i=1}^N f_i(x) of their individual convex costs. Existing literature proposes distributed gradient-like methods that are computationally cheap and resilient to link failures, but have slow convergence rates. In this paper, we propose accelerated distributed gradient methods that: 1) are resilient to link failures; 2) computationally cheap; and 3) improve convergence rates over other gradient methods. We model the network by a sequence of independent, identically distributed random matrices {W(k)} drawn from the set of symmetric, stochastic matrices with positive diagonals. The network is connected on average and the cost functions are convex, differentiable, with Lipschitz continuous and bounded gradients. We design two distributed Nesterov-like gradient methods that modify the D-NG and D-NC methods that we proposed for static networks. We prove their convergence rates in terms of the expected optimality gap at the cost function. Let k and K be the number of per-node gradient evaluations and per-node communications, respectively. Then the modified D-NG achieves rates O(log k/k) and O(\log K/K), and the modified D-NC rates O(1/k^2) and O(1/K^{2-ξ}), where ξ>0 is arbitrarily small. For comparison, the standard distributed gradient method cannot do better than Ω(1/k^{2/3}) and Ω(1/K^{2/3}), on the same class of cost functions (even for static networks). Simulation examples illustrate our analytical findings.
journal; submitted for publication on May 11, 2013
Cited by in corpus (13)
- Asynchronous Distributed Optimization over Lossy Networks via Relaxed ADMM: Stability and Linear Convergence
- Distributed Online Optimization in Time-Varying Unbalanced Networks without Explicit Subgradients
- Distributed Constrained Recursive Nonlinear Least-Squares Estimation: Algorithms and Asymptotics
- Accelerated Distributed Dual Averaging over Evolving Networks of Growing Connectivity
- Communication-Efficient Distributed Strongly Convex Stochastic Optimization: Non-Asymptotic Rates
- Accelerated Gradient Tracking over Time-varying Graphs for Decentralized Optimization
- Optimization over time-varying directed graphs with row and column-stochastic matrices
- Rapid Transitions with Robust Accelerated Delayed Self Reinforcement for Consensus-Based Networks
- A Unification and Generalization of Exact Distributed First Order Methods
- A general framework for decentralized optimization with first-order methods
- A New Approach for Optimizing Highly Nonlinear Problems Based on the Observer Effect Concept
- Modified swarm-based metaheuristics enhance Gradient Descent initialization performance: Application for EEG spatial filtering
- Decentralized Composite Optimization in Stochastic Networks: A Dual Averaging Approach with Linear Convergence