ProgFed: Effective, Communication, and Computation Efficient Federated Learning by Progressive Training
arXiv:2110.05323
Abstract
Federated learning is a powerful distributed learning scheme that allows numerous edge devices to collaboratively train a model without sharing their data. However, training is resource-intensive for edge devices, and limited network bandwidth is often the main bottleneck. Prior work often overcomes the constraints by condensing the models or messages into compact formats, e.g., by gradient compression or distillation. In contrast, we propose ProgFed, the first progressive training framework for efficient and effective federated learning. It inherently reduces computation and two-way communication costs while maintaining the strong performance of the final models. We theoretically prove that ProgFed converges at the same asymptotic rate as standard training on full models. Extensive results on a broad range of architectures, including CNNs (VGG, ResNet, ConvNets) and U-nets, and diverse tasks from simple classification to medical image segmentation show that our highly effective training approach saves up to computation and up to communication costs for converged models. As our approach is also complimentary to prior work on compression, we can achieve a wide range of trade-offs by combining these techniques, showing reduced communication of up to at only loss in utility. Code is available at https://github.com/hui-po-wang/ProgFed.
To appear in ICML 2022
References in corpus (15)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Distilling the Knowledge in a Neural Network
- Federated Learning: Strategies for Improving Communication Efficiency
- Improved Regularization of Convolutional Neural Networks with Cutout
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- SGDR: Stochastic Gradient Descent with Warm Restarts
- Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training
- TernGrad: Ternary Gradients to Reduce Communication in Distributed Deep Learning
- Ensemble Distillation for Robust Model Fusion in Federated Learning
- FedMD: Heterogenous Federated Learning via Model Distillation
- Model compression via distillation and quantization
- Slimmable Neural Networks
- DoubleSqueeze: Parallel Stochastic Gradient Descent with Double-Pass Error-Compensated Compression
- Dynamic Model Pruning with Feedback
- On Communication Compression for Distributed Optimization on Heterogeneous Data