Synthetic data shuffling accelerates the convergence of federated learning under data heterogeneity
arXiv:2306.13263
Abstract
In federated learning, data heterogeneity is a critical challenge. A straightforward solution is to shuffle the clients' data to homogenize the distribution. However, this may violate data access rights, and how and when shuffling can accelerate the convergence of a federated optimization algorithm is not theoretically well understood. In this paper, we establish a precise and quantifiable correspondence between data heterogeneity and parameters in the convergence rate when a fraction of data is shuffled across clients. We prove that shuffling can quadratically reduce the gradient dissimilarity with respect to the shuffling percentage, accelerating convergence. Inspired by the theory, we propose a practical approach that addresses the data access rights issue by shuffling locally generated synthetic data. The experimental results show that shuffling synthetic data improves the performance of multiple existing federated learning algorithms by a large margin.
Accepted at TMLR
References in corpus (14)
- Federated Optimization: Distributed Machine Learning for On-Device Intelligence
- Tackling the Objective Inconsistency Problem in Heterogeneous Federated Optimization
- A Field Guide to Federated Optimization
- Mime: Mimicking Centralized Stochastic Algorithms in Federated Learning
- FedMix: Approximation of Mixup under Mean Augmented Federated Learning
- FedAvg with Fine Tuning: Local Updates Lead to Representation Learning
- Linear Convergence in Federated Learning: Tackling Client Heterogeneity and Sparse Gradients
- Federated Learning via Synthetic Data
- DENSE: Data-Free One-Shot Federated Learning
- Differentially Private Diffusion Models
- On the Unreasonable Effectiveness of Federated Averaging with Heterogeneous Data
- TCT: Convexifying Federated Learning using Bootstrapped Neural Tangent Kernels
- On the effectiveness of partial variance reduction in federated learning with heterogeneous data
- FedLAP-DP: Federated Learning by Sharing Differentially Private Loss Approximations