Few-shot learning of neural networks from scratch by pseudo example optimization
arXiv:1802.03039
Abstract
In this paper, we propose a simple but effective method for training neural networks with a limited amount of training data. Our approach inherits the idea of knowledge distillation that transfers knowledge from a deep or wide reference model to a shallow or narrow target model. The proposed method employs this idea to mimic predictions of reference estimators that are more robust against overfitting than the network we want to train. Different from almost all the previous work for knowledge distillation that requires a large amount of labeled training data, the proposed method requires only a small amount of training data. Instead, we introduce pseudo training examples that are optimized as a part of model parameters. Experimental results for several benchmark datasets demonstrate that the proposed method outperformed all the other baselines, such as naive training of the target model and standard knowledge distillation.
14 pages, 2 figures, will be presented at BMVC2018
Cited by in corpus (9)
- Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
- Knowledge Distillation in Deep Learning and its Applications
- Dreaming to Distill: Data-free Knowledge Transfer via DeepInversion
- Zero-shot Knowledge Transfer via Adversarial Belief Matching
- Few Sample Knowledge Distillation for Efficient Network Compression
- Neural Networks Are More Productive Teachers Than Human Raters: Active Mixup for Data-Efficient Knowledge Distillation from a Blackbox Model
- Confidence Conditioned Knowledge Distillation
- Knowledge Distillation By Sparse Representation Matching
- On Cross-Layer Alignment for Model Fusion of Heterogeneous Neural Networks