Rethinking Architecture Design for Tackling Data Heterogeneity in Federated Learning
arXiv:2106.06047
Abstract
Federated learning is an emerging research paradigm enabling collaborative training of machine learning models among different organizations while keeping data private at each institution. Despite recent progress, there remain fundamental challenges such as the lack of convergence and the potential for catastrophic forgetting across real-world heterogeneous devices. In this paper, we demonstrate that self-attention-based architectures (e.g., Transformers) are more robust to distribution shifts and hence improve federated learning over heterogeneous data. Concretely, we conduct the first rigorous empirical investigation of different neural architectures across a range of federated algorithms, real-world benchmarks, and heterogeneous data splits. Our experiments show that simply replacing convolutional networks with Transformers can greatly reduce catastrophic forgetting of previous devices, accelerate convergence, and reach a better global model, especially when dealing with heterogeneous data. We release our code and pretrained models at https://github.com/Liangqiong/ViT-FL-main to encourage future exploration in robust architectures as an alternative to current research efforts on the optimization front.
Published as a conference paper at CVPR 2022
References in corpus (10)
- Measuring the Effects of Non-Identical Data Distribution for Federated Visual Classification
- Federated Learning with Personalization Layers
- Ensemble Distillation for Robust Model Fusion in Federated Learning
- Generating Long Sequences with Sparse Transformers
- Inverting Gradients -- How easy is it to break privacy in federated learning?
- Measuring the tendency of CNNs to Learn Surface Statistical Regularities
- Differential Privacy-enabled Federated Learning for Sensitive Health Data
- Think Locally, Act Globally: Federated Learning with Local and Global Representations
- Pretrained Transformers as Universal Computation Engines
- An Embarrassingly Simple Approach for Transfer Learning from Pretrained Language Models