Training Shallow and Thin Networks for Acceleration via Knowledge Distillation with Conditional Adversarial Networks
arXiv:1709.00513
Abstract
There is an increasing interest on accelerating neural networks for real-time applications. We study the student-teacher strategy, in which a small and fast student network is trained with the auxiliary information learned from a large and accurate teacher network. We propose to use conditional adversarial networks to learn the loss function to transfer knowledge from teacher to student. The proposed method is particularly effective for relatively small student networks. Moreover, experimental results show the effect of network size when the modern networks are used as student. We empirically study the trade-off between inference time and classification accuracy, and provide suggestions on choosing a proper student network.
Shorter version will appear at ICLR workshop 2018
References in corpus (10)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Distilling the Knowledge in a Neural Network
- Conditional Generative Adversarial Nets
- Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer
- A Downsampled Variant of ImageNet as an Alternative to the CIFAR datasets
- A Closer Look at Memorization in Deep Networks
- Like What You Like: Knowledge Distill via Neuron Selectivity Transfer
- Distral: Robust Multitask Reinforcement Learning
- Why Deep Neural Networks for Function Approximation?
- Deep Model Compression: Distilling Knowledge from Noisy Teachers
Cited by in corpus (11)
- Knowledge Distillation: A Survey
- Cross-Modality Knowledge Distillation Network for Monocular 3D Object Detection
- Compacting Deep Neural Networks for Internet of Things: Methods and Applications
- Distilling and Transferring Knowledge via cGAN-generated Samples for Image Classification and Regression
- Preparing Lessons: Improve Knowledge Distillation with Better Supervision
- Online Knowledge Distillation via Multi-branch Diversity Enhancement
- Improving Face Recognition from Hard Samples via Distribution Distillation Loss
- Prune Your Model Before Distill It
- Complementary Relation Contrastive Distillation
- CoCo DistillNet: a Cross-layer Correlation Distillation Network for Pathological Gastric Cancer Segmentation
- Private Knowledge Transfer via Model Distillation with Generative Adversarial Networks