Federated Learning Based on Dynamic Regularization
arXiv:2111.04263
Abstract
We propose a novel federated learning method for distributively training neural network models, where the server orchestrates cooperation between a subset of randomly chosen devices in each round. We view Federated Learning problem primarily from a communication perspective and allow more device level computations to save transmission costs. We point out a fundamental dilemma, in that the minima of the local-device level empirical loss are inconsistent with those of the global empirical loss. Different from recent prior works, that either attempt inexact minimization or utilize devices for parallelizing gradient computation, we propose a dynamic regularizer for each device at each round, so that in the limit the global and device solutions are aligned. We demonstrate both through empirical results on real and synthetic data as well as analytical results that our scheme leads to efficient training, in both convex and non-convex settings, while being fully agnostic to device heterogeneity and robust to large number of devices, partial participation and unbalanced data.
Slightly extended version of ICLR 2021 Paper
References in corpus (5)
- Federated Optimization: Distributed Machine Learning for On-Device Intelligence
- Bayesian Nonparametric Federated Learning of Neural Networks
- Variance Reduced Local SGD with Lower Communication Complexity
- FedSplit: An algorithmic framework for fast federated optimization
- A Unified Analysis of Stochastic Gradient Methods for Nonconvex Federated Optimization
Cited by in corpus (6)
- A Field Guide to Federated Optimization
- No Fear of Heterogeneity: Classifier Calibration for Federated Learning with Non-IID Data
- Local-Global Knowledge Distillation in Heterogeneous Federated Learning with Non-IID Data
- Model-Contrastive Federated Learning
- Federated Hyperparameter Tuning: Challenges, Baselines, and Connections to Weight-Sharing
- Distantly Supervised Relation Extraction in Federated Settings