Mime: Mimicking Centralized Stochastic Algorithms in Federated Learning
arXiv:2008.03606
Abstract
Federated learning (FL) is a challenging setting for optimization due to the heterogeneity of the data across different clients which gives rise to the client drift phenomenon. In fact, obtaining an algorithm for FL which is uniformly better than simple centralized training has been a major open problem thus far. In this work, we propose a general algorithmic framework, Mime, which i) mitigates client drift and ii) adapts arbitrary centralized optimization algorithms such as momentum and Adam to the cross-device federated learning setting. Mime uses a combination of control-variates and server-level statistics (e.g. momentum) at every client-update step to ensure that each local update mimics that of the centralized method run on iid data. We prove a reduction result showing that Mime can translate the convergence of a generic algorithm in the centralized setting into convergence in the federated setting. Further, we show that when combined with momentum based variance reduction, Mime is provably faster than any centralized method--the first such result. We also perform a thorough experimental exploration of Mime's performance on real world datasets.
Version 2 provides stronger theoretical results and more thorough experiments
References in corpus (18)
- Federated Optimization: Distributed Machine Learning for On-Device Intelligence
- On the Convergence of Adam and Beyond
- Towards Federated Learning at Scale: System Design
- Large Batch Training of Convolutional Networks
- Agnostic Federated Learning
- A Unified Theory of Decentralized SGD with Changing Topology and Local Updates
- Expanding the Reach of Federated Learning by Reducing Client Resource Requirements
- Error Feedback Fixes SignSGD and other Gradient Compression Schemes
- On the Linear Speedup Analysis of Communication Efficient Momentum SGD for Distributed Non-Convex Optimization
- Adaptive Federated Optimization
- Federated Learning Based on Dynamic Regularization
- Attack of the Tails: Yes, You Really Can Backdoor Federated Learning
- AIDE: Fast and Communication Efficient Distributed Optimization
- Minibatch vs Local SGD for Heterogeneous Distributed Learning
- Secure Byzantine-Robust Machine Learning
- On the Outsized Importance of Learning Rates in Local Update Methods
- Understanding Unintended Memorization in Federated Learning
- Communication trade-offs for synchronized distributed SGD with large step size
Cited by in corpus (32)
- A Field Guide to Federated Optimization
- FedCluster: Boosting the Convergence of Federated Learning via Cluster-Cycling
- Federated Learning with Nesterov Accelerated Gradient
- FedLGA: Towards System-Heterogeneity of Federated Learning via Local Gradient Approximation
- Federated Learning via Posterior Averaging: A New Perspective and Practical Algorithms
- Quasi-Global Momentum: Accelerating Decentralized Deep Learning on Heterogeneous Data
- Differentially Private Federated Learning on Heterogeneous Data
- Learning from History for Byzantine Robust Optimization
- FedCM: Federated Learning with Client-level Momentum
- Local Adaptivity in Federated Learning: Convergence and Consistency
- STEM: A Stochastic Two-Sided Momentum Algorithm Achieving Near-Optimal Sample and Communication Complexities for Federated Learning
- FedJAX: Federated learning simulation with JAX
- Bias-Variance Reduced Local SGD for Less Heterogeneous Federated Learning
- FedDR -- Randomized Douglas-Rachford Splitting Algorithms for Nonconvex Federated Composite Optimization
- Faster Non-Convex Federated Learning via Global and Local Momentum
- Federated Face Recognition
- Efficient Algorithms for Federated Saddle Point Optimization
- Hybrid Federated Learning: Algorithms and Implementation
- Anarchic Federated Learning
- Jointly Learning from Decentralized (Federated) and Centralized Data to Mitigate Distribution Shift
- Learning Federated Representations and Recommendations with Limited Negatives
- Synthetic data shuffling accelerates the convergence of federated learning under data heterogeneity
- Towards Model Agnostic Federated Learning Using Knowledge Distillation
- On Large-Cohort Training for Federated Learning
- Federated Functional Gradient Boosting
- Communication-Efficient Agnostic Federated Averaging
- Linear Speedup in Personalized Collaborative Learning
- Multi-task Federated Edge Learning (MtFEEL) in Wireless Networks
- Compositional federated learning: Applications in distributionally robust averaging and meta learning
- TOFU: Towards Obfuscated Federated Updates by Encoding Weight Updates into Gradients from Proxy Data
- On Second-order Optimization Methods for Federated Learning
- Towards Heterogeneous Clients with Elastic Federated Learning