FedCluster: Boosting the Convergence of Federated Learning via Cluster-Cycling
arXiv:2009.10748 · doi:10.1109/BigData50022.2020.9377960
Abstract
We develop FedCluster--a novel federated learning framework with improved optimization efficiency, and investigate its theoretical convergence properties. The FedCluster groups the devices into multiple clusters that perform federated learning cyclically in each learning round. Therefore, each learning round of FedCluster consists of multiple cycles of meta-update that boost the overall convergence. In nonconvex optimization, we show that FedCluster with the devices implementing the local {stochastic gradient descent (SGD)} algorithm achieves a faster convergence rate than the conventional {federated averaging (FedAvg)} algorithm in the presence of device-level data heterogeneity. We conduct experiments on deep learning applications and demonstrate that FedCluster converges significantly faster than the conventional federated learning under diverse levels of device-level data heterogeneity for a variety of local optimizers.
10 pages, 6 figures
References in corpus (16)
- Federated Optimization: Distributed Machine Learning for On-Device Intelligence
- Tackling the Objective Inconsistency Problem in Heterogeneous Federated Optimization
- Adaptive Personalized Federated Learning
- Agnostic Federated Learning
- Federated Learning of a Mixture of Global and Local Models
- A Unified Theory of Decentralized SGD with Changing Topology and Local Updates
- Adaptive Federated Optimization
- A Framework for Evaluating Gradient Leakage Attacks in Federated Learning
- Mime: Mimicking Centralized Stochastic Algorithms in Federated Learning
- Variance Reduced Local SGD with Lower Communication Complexity
- Minibatch vs Local SGD for Heterogeneous Distributed Learning
- Federated Residual Learning
- FedMAX: Mitigating Activation Divergence for Accurate and Communication-Efficient Federated Learning
- Privacy-Preserving Blockchain Based Federated Learning with Differential Data Sharing
- Adaptive Sampling Distributed Stochastic Variance Reduced Gradient for Heterogeneous Distributed Datasets
- Multi-Level Local SGD for Heterogeneous Hierarchical Networks
Cited by in corpus (6)
- Edge Learning for B5G Networks with Distributed Signal Processing: Semantic Communication, Edge Computing, and Wireless Sensing
- FedCross: Towards Accurate Federated Learning via Multi-Model Cross-Aggregation
- Is Aggregation the Only Choice? Federated Learning via Layer-wise Model Recombination
- Federated Learning for Commercial Image Sources
- Certifiably-Robust Federated Adversarial Learning via Randomized Smoothing
- Research on Resource Allocation for Efficient Federated Learning