Knowledge distillation using unlabeled mismatched images
arXiv:1703.07131
Abstract
Current approaches for Knowledge Distillation (KD) either directly use training data or sample from the training data distribution. In this paper, we demonstrate effectiveness of 'mismatched' unlabeled stimulus to perform KD for image classification networks. For illustration, we consider scenarios where this is a complete absence of training data, or mismatched stimulus has to be used for augmenting a small amount of training data. We demonstrate that stimulus complexity is a key factor for distillation's good performance. Our examples include use of various datasets for stimulating MNIST and CIFAR teachers.
References in corpus (2)
Cited by in corpus (5)
- Dream Distillation: A Data-Independent Model Compression Framework
- Mosaicking to Distill: Knowledge Distillation from Out-of-Domain Data
- Knowledge distillation for optimization of quantized deep neural networks
- EdgeAI: A Vision for Deep Learning in IoT Era
- New Directions in Distributed Deep Learning: Bringing the Network at Forefront of IoT Design