Unifying distillation and privileged information
arXiv:1511.03643
Abstract
Distillation (Hinton et al., 2015) and privileged information (Vapnik & Izmailov, 2015) are two techniques that enable machines to learn from other machines. This paper unifies these two techniques into generalized distillation, a framework to learn from multiple machines and data representations. We provide theoretical and causal insight about the inner workings of generalized distillation, extend it to unsupervised, semisupervised and multitask learning scenarios, and illustrate its efficacy on a variety of numerical simulations on both synthetic and real-world data.
References in corpus (3)
Cited by in corpus (38)
- Knowledge Distillation: A Survey
- Distilled Siamese Networks for Visual Tracking
- Understanding and Improving Knowledge Distillation
- Learning with privileged information via adversarial discriminative modality distillation
- Learning from Noisy Labels with Distillation
- Compacting Deep Neural Networks for Internet of Things: Methods and Applications
- Synergic Adversarial Label Learning for Grading Retinal Diseases via Knowledge Distillation and Multi-task Learning
- Data Distillation: Towards Omni-Supervised Learning
- Hydra: Preserving Ensemble Diversity for Model Distillation
- Spatio-Temporal Graph for Video Captioning with Knowledge Distillation
- SPIGAN: Privileged Adversarial Learning from Simulation
- Smooth Neighbors on Teacher Graphs for Semi-supervised Learning
- Adaptive Regularization of Labels
- Active Long Term Memory Networks
- Assisted Learning: A Framework for Multi-Organization Learning
- OMNIA Faster R-CNN: Detection in the wild through dataset merging and soft distillation
- DMCL: Distillation Multiple Choice Learning for Multimodal Action Recognition
- In Defense of the Triplet Loss Again: Learning Robust Person Re-Identification with Fast Approximated Triplet Loss and Label Distillation
- GAN Slimming: All-in-One GAN Compression by A Unified Optimization Framework
- Flow-Distilled IP Two-Stream Networks for Compressed Video Action Recognition
- Privileged Information Dropout in Reinforcement Learning
- Distilling EEG Representations via Capsules for Affective Computing
- Ego-Exo: Transferring Visual Representations from Third-person to First-person Videos
- Hard Pixel Mining for Depth Privileged Semantic Segmentation
- Accelerating Large Scale Knowledge Distillation via Dynamic Importance Sampling
- SECS: Efficient Deep Stream Processing via Class Skew Dichotomy
- Semi-Supervised 3D Hand-Object Poses Estimation with Interactions in Time
- Co-advise: Cross Inductive Bias Distillation
- Embracing the Dark Knowledge: Domain Generalization Using Regularized Knowledge Distillation
- Churn Reduction via Distillation
- Pro-KD: Progressive Distillation by Following the Footsteps of the Teacher
- Improving the trustworthiness of image classification models by utilizing bounding-box annotations
- Improving Span-based Question Answering Systems with Coarsely Labeled Data
- Defocus Blur Detection via Depth Distillation
- An Exploration of Mimic Architectures for Residual Network Based Spectral Mapping
- MIML-FCN+: Multi-instance Multi-label Learning via Fully Convolutional Networks with Privileged Information
- Private Knowledge Transfer via Model Distillation with Generative Adversarial Networks
- BridgeNets: Student-Teacher Transfer Learning Based on Recursive Neural Networks and its Application to Distant Speech Recognition