FedKD: Communication Efficient Federated Learning via Knowledge Distillation
arXiv:2108.13323 · doi:10.1038/s41467-022-29763-x
Abstract
Federated learning is widely used to learn intelligent models from decentralized data. In federated learning, clients need to communicate their local model updates in each iteration of model learning. However, model updates are large in size if the model contains numerous parameters, and there usually needs many rounds of communication until model converges. Thus, the communication cost in federated learning can be quite heavy. In this paper, we propose a communication efficient federated learning method based on knowledge distillation. Instead of directly communicating the large models between clients and server, we propose an adaptive mutual distillation framework to reciprocally learn a student and a teacher model on each client, where only the student model is shared by different clients and updated collaboratively to reduce the communication cost. Both the teacher and student on each client are learned on its local data and the knowledge distilled from each other, where their distillation intensities are controlled by their prediction quality. To further reduce the communication cost, we propose a dynamic gradient approximation method based on singular value decomposition to approximate the exchanged gradients with dynamic precision. Extensive experiments on benchmark datasets in different tasks show that our approach can effectively reduce the communication cost and achieve competitive results.
References in corpus (9)
- Distilling the Knowledge in a Neural Network
- FedKD: Communication Efficient Federated Learning via Knowledge Distillation
- Federated Machine Learning: Concept and Applications
- FedMD: Heterogenous Federated Learning via Model Distillation
- Distilling Task-Specific Knowledge from BERT into Simple Neural Networks
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training
- DKN: Deep Knowledge-Aware Network for News Recommendation
- Fast Federated Learning by Balancing Communication Trade-Offs
- Federated Knowledge Distillation
Cited by in corpus (24)
- FedKD: Communication Efficient Federated Learning via Knowledge Distillation
- Decentralized Federated Learning: Fundamentals, State of the Art, Frameworks, Trends, and Challenges
- Advances of Machine Learning in Materials Science: Ideas and Techniques
- Explainable, Domain-Adaptive, and Federated Artificial Intelligence in Medicine
- FedICT: Federated Multi-task Distillation for Multi-access Edge Computing
- Advances in Robust Federated Learning: A Survey with Heterogeneity Considerations
- Model Pruning Enables Localized and Efficient Federated Learning for Yield Forecasting and Data Sharing
- Seamless Detection: Unifying Salient Object Detection and Camouflaged Object Detection
- A Comprehensive Review and a Taxonomy of Edge Machine Learning: Requirements, Paradigms, and Techniques
- A Privacy Preserving System for Movie Recommendations Using Federated Learning
- Communication-Efficient Training Workload Balancing for Decentralized Multi-Agent Learning
- Knowledge Distillation in Federated Learning: a Survey on Long Lasting Challenges and New Solutions
- Foundational Models and Federated Learning: Survey, Taxonomy, Challenges and Practical Insights
- FLEdge: Benchmarking Federated Machine Learning Applications in Edge Computing Systems
- Cyber Attacks Prevention Towards Prosumer-based EV Charging Stations: An Edge-assisted Federated Prototype Knowledge Distillation Approach
- Rethinking Knowledge Distillation in Collaborative Machine Learning: Memory, Knowledge, and Their Interactions
- UNIDEAL: Curriculum Knowledge Distillation Federated Learning
- Explainable Semantic Federated Learning Enabled Industrial Edge Network for Fire Surveillance
- Decentralized Personalization for Federated Medical Image Segmentation via Gossip Contrastive Mutual Learning
- FedBrain-Distill: Communication-Efficient Federated Brain Tumor Classification Using Ensemble Knowledge Distillation on Non-IID Data
- FedMoE-DA: Federated Mixture of Experts via Domain Aware Fine-grained Aggregation
- FedAGHN: Personalized Federated Learning with Attentive Graph HyperNetworks
- Communication-Aware Knowledge Distillation for Federated LLM Fine-Tuning over Wireless Networks
- Security and Privacy Issues and Solutions in Federated Learning for Digital Healthcare