Rethinking Knowledge Distillation in Collaborative Machine Learning: Memory, Knowledge, and Their Interactions
arXiv:2512.19972 · doi:10.1109/TNSE.2025.3572362
Abstract
Collaborative learning has emerged as a key paradigm in large-scale intelligent systems, enabling distributed agents to cooperatively train their models while addressing their privacy concerns. Central to this paradigm is knowledge distillation (KD), a technique that facilitates efficient knowledge transfer among agents. However, the underlying mechanisms by which KD leverages memory and knowledge across agents remain underexplored. This paper aims to bridge this gap by offering a comprehensive review of KD in collaborative learning, with a focus on the roles of memory and knowledge. We define and categorize memory and knowledge within the KD process and explore their interrelationships, providing a clear understanding of how knowledge is extracted, stored, and shared in collaborative settings. We examine various collaborative learning patterns, including distributed, hierarchical, and decentralized structures, and provide insights into how memory and knowledge dynamics shape the effectiveness of KD in collaborative learning. Particularly, we emphasize task heterogeneity in distributed learning pattern covering federated learning (FL), multi-agent domain adaptation (MADA), federated multi-modal learning (FML), federated continual learning (FCL), federated multi-task learning (FMTL), and federated graph knowledge embedding (FKGE). Additionally, we highlight model heterogeneity, data heterogeneity, resource heterogeneity, and privacy concerns of these tasks. Our analysis categorizes existing work based on how they handle memory and knowledge. Finally, we discuss existing challenges and propose future directions for advancing KD techniques in the context of collaborative learning.
Published in IEEE TNSE
References in corpus (21)
- Overcoming catastrophic forgetting in neural networks
- Knowledge Distillation: A Survey
- Riemannian Walk for Incremental Learning: Understanding Forgetting and Intransigence
- Knowledge Distillation and Student-Teacher Learning for Visual Intelligence: A Review and New Outlooks
- FedKD: Communication Efficient Federated Learning via Knowledge Distillation
- Low-resolution Face Recognition in the Wild via Selective Knowledge Distillation
- Explaining Neural Scaling Laws
- Federated Continual Learning via Knowledge Fusion: A Survey
- DE-RRD: A Knowledge Distillation Framework for Recommender System
- On Representation Knowledge Distillation for Graph Neural Networks
- FedICT: Federated Multi-task Distillation for Multi-access Edge Computing
- A Survey on Symbolic Knowledge Distillation of Large Language Models
- MulDE: Multi-teacher Knowledge Distillation for Low-dimensional Knowledge Graph Embeddings
- Distilling Knowledge by Mimicking Features
- Knowledge Distillation approach towards Melanoma Detection
- Adversarial Feature Alignment: Avoid Catastrophic Forgetting in Incremental Task Lifelong Learning
- Contextual Distillation Model for Diversified Recommendation
- Progressive Distillation Based on Masked Generation Feature Method for Knowledge Graph Completion
- CrowdTransfer: Enabling Crowd Knowledge Transfer in AIoT Community
- Multi-source-free Domain Adaptation via Uncertainty-aware Adaptive Distillation
- : Language Modeling with Explicit Memory