FLIX: A Simple and Communication-Efficient Alternative to Local Methods in Federated Learning
arXiv:2111.11556
Abstract
Federated Learning (FL) is an increasingly popular machine learning paradigm in which multiple nodes try to collaboratively learn under privacy, communication and multiple heterogeneity constraints. A persistent problem in federated learning is that it is not clear what the optimization objective should be: the standard average risk minimization of supervised learning is inadequate in handling several major constraints specific to federated learning, such as communication adaptivity and personalization control. We identify several key desiderata in frameworks for federated learning and introduce a new framework, FLIX, that takes into account the unique challenges brought by federated learning. FLIX has a standard finite-sum form, which enables practitioners to tap into the immense wealth of existing (potentially non-local) methods for distributed optimization. Through a smart initialization that does not require any communication, FLIX does not require the use of local steps but is still provably capable of performing dissimilarity regularization on par with local methods. We give several algorithms for solving the FLIX formulation efficiently under communication constraints. Finally, we corroborate our theoretical results with extensive experimentation.
V2: includes non-convex analysis as well as new large-scale experiments with neural networks. To appear in AISTATS 2022
References in corpus (17)
- Improving Federated Learning Personalization via Model Agnostic Meta Learning
- Adaptive Personalized Federated Learning
- Three Approaches for Personalization with Applications to Federated Learning
- Federated Learning of a Mixture of Global and Local Models
- Better Mini-Batch Algorithms via Accelerated Gradient Methods
- One-Shot Federated Learning
- FedSplit: An algorithmic framework for fast federated optimization
- Salvaging Federated Learning by Local Adaptation
- Better Theory for SGD in the Nonconvex World
- Minibatch vs Local SGD for Heterogeneous Distributed Learning
- Personalized Cross-Silo Federated Learning on Non-IID Data
- A Unified Analysis of Stochastic Gradient Methods for Nonconvex Federated Optimization
- Specialized federated learning using a mixture of experts
- Minimax Estimation for Personalized Federated Learning: An Alternative between FedAvg and Local Training?
- Modeling and Optimization Trade-off in Meta-learning
- How Fine-Tuning Allows for Effective Meta-Learning
- More Industry-friendly: Federated Learning with High Efficient Design