FedICT: Federated Multi-task Distillation for Multi-access Edge Computing
arXiv:2301.00389 · doi:10.1109/TPDS.2023.3289444
Abstract
The growing interest in intelligent services and privacy protection for mobile devices has given rise to the widespread application of federated learning in Multi-access Edge Computing (MEC). Diverse user behaviors call for personalized services with heterogeneous Machine Learning (ML) models on different devices. Federated Multi-task Learning (FMTL) is proposed to train related but personalized ML models for different devices, whereas previous works suffer from excessive communication overhead during training and neglect the model heterogeneity among devices in MEC. Introducing knowledge distillation into FMTL can simultaneously enable efficient communication and model heterogeneity among clients, whereas existing methods rely on a public dataset, which is impractical in reality. To tackle this dilemma, Federated MultI-task Distillation for Multi-access Edge CompuTing (FedICT) is proposed. FedICT direct local-global knowledge aloof during bi-directional distillation processes between clients and the server, aiming to enable multi-task clients while alleviating client drift derived from divergent optimization directions of client-side local models. Specifically, FedICT includes Federated Prior Knowledge Distillation (FPKD) and Local Knowledge Adjustment (LKA). FPKD is proposed to reinforce the clients' fitting of local data by introducing prior knowledge of local data distributions. Moreover, LKA is proposed to correct the distillation loss of the server, making the transferred local knowledge better match the generalized representation. Experiments on three datasets show that FedICT significantly outperforms all compared benchmarks in various data heterogeneous and model architecture settings, achieving improved accuracy with less than 1.2% training communication overhead compared with FedAvg and no more than 75% training communication round compared with FedGKT.
Accepted by IEEE TRANSACTIONS ON PARALLEL AND DISTRIBUTED SYSTEMS
References in corpus (15)
- Distilling the Knowledge in a Neural Network
- Towards Personalized Federated Learning
- Knowledge Distillation and Student-Teacher Learning for Visual Intelligence: A Review and New Outlooks
- FedKD: Communication Efficient Federated Learning via Knowledge Distillation
- FedMD: Heterogenous Federated Learning via Model Distillation
- Communication-Efficient On-Device Machine Learning: Federated Distillation and Augmentation under Non-IID Private Data
- FedML: A Research Library and Benchmark for Federated Machine Learning
- Distillation-Based Semi-Supervised Federated Learning for Communication-Efficient Collaborative Training with Non-IID Private Data
- CINIC-10 is not ImageNet or CIFAR-10
- Cronus: Robust and Heterogeneous Collaborative Learning with Black-Box Knowledge Transfer
- Local-Global Knowledge Distillation in Heterogeneous Federated Learning with Non-IID Data
- FedGEMS: Federated Learning of Larger Server Models via Selective Knowledge Fusion
- Personalized Federated Learning for Heterogeneous Clients with Clustered Knowledge Transfer
- FedDTG:Federated Data-Free Knowledge Distillation via Three-Player Generative Adversarial Networks
- Spirit Distillation: Precise Real-time Semantic Segmentation of Road Scenes with Insufficient Data
Cited by in corpus (3)
- Emerging Trends in Federated Learning: From Model Fusion to Federated X Learning
- Rethinking Knowledge Distillation in Collaborative Machine Learning: Memory, Knowledge, and Their Interactions
- Multi-task Federated Learning with Encoder-Decoder Structure: Enabling Collaborative Learning Across Different Tasks