Deep Mutual Learning
arXiv:1706.00384
Abstract
Model distillation is an effective and widely used technique to transfer knowledge from a teacher to a student network. The typical application is to transfer from a powerful large network or ensemble to a small network, that is better suited to low-memory or fast execution requirements. In this paper, we present a deep mutual learning (DML) strategy where, rather than one way transfer between a static pre-defined teacher and a student, an ensemble of students learn collaboratively and teach each other throughout the training process. Our experiments show that a variety of network architectures benefit from mutual learning and achieve compelling results on CIFAR-100 recognition and Market-1501 person re-identification benchmarks. Surprisingly, it is revealed that no prior powerful teacher network is necessary -- mutual learning of a collection of simple student networks works, and moreover outperforms distillation from a more powerful yet static teacher.
10 pages, 4 figures
References in corpus (1)
Cited by in corpus (13)
- AlignedReID: Surpassing Human-Level Performance in Person Re-Identification
- Margin Sample Mining Loss: A Deep Learning Based Method for Person Re-identification
- Re-ID done right: towards good practices for person re-identification
- Distilling Policy Distillation
- Knowledge Distillation via Route Constrained Optimization
- FEED: Feature-level Ensemble for Knowledge Distillation
- Intra-Inter Camera Similarity for Unsupervised Person Re-Identification
- Towards Cross-modality Medical Image Segmentation with Online Mutual Knowledge Distillation
- Backbone Can Not be Trained at Once: Rolling Back to Pre-trained Network for Person Re-Identification
- Simple Distillation Baselines for Improving Small Self-supervised Models
- MOD: A Deep Mixture Model with Online Knowledge Distillation for Large Scale Video Temporal Concept Localization
- On the Orthogonality of Knowledge Distillation with Other Techniques: From an Ensemble Perspective
- Improving Route Choice Models by Incorporating Contextual Factors via Knowledge Distillation