Knowledge Distillation in Deep Learning and its Applications
arXiv:2007.09029 · doi:10.7717/peerj-cs.474
Abstract
Deep learning based models are relatively large, and it is hard to deploy such models on resource-limited devices such as mobile phones and embedded devices. One possible solution is knowledge distillation whereby a smaller model (student model) is trained by utilizing the information from a larger model (teacher model). In this paper, we present a survey of knowledge distillation techniques applied to deep learning models. To compare the performances of different techniques, we propose a new metric called distillation metric. Distillation metric compares different knowledge distillation algorithms based on sizes and accuracy scores. Based on the survey, some interesting conclusions are drawn and presented in this paper.
9 pages
References in corpus (7)
- Distilling the Knowledge in a Neural Network
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Learning Face Representation from Scratch
- Data-Free Knowledge Distillation for Deep Neural Networks
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning
- Zero-Shot Knowledge Distillation in Deep Networks