Cross Modal Distillation for Supervision Transfer
arXiv:1507.00448
Abstract
In this work we propose a technique that transfers supervision between images from different modalities. We use learned representations from a large labeled modality as a supervisory signal for training representations for a new unlabeled paired modality. Our method enables learning of rich representations for unlabeled modalities and can be used as a pre-training procedure for new modalities with limited labeled data. We show experimental results where we transfer supervision from labeled RGB images to unlabeled depth and optical flow images and demonstrate large improvements for both these cross modal supervision transfers. Code, data and pre-trained models are available at https://github.com/s-gupta/fast-rcnn/tree/distillation
Updated version (v2) contains additional experiments and results
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Distilling the Knowledge in a Neural Network
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Learning Transferable Features with Deep Adaptation Networks
- Deep Domain Confusion: Maximizing for Domain Invariance
- FitNets: Hints for Thin Deep Nets
- Return of the Devil in the Details: Delving Deep into Convolutional Nets
- Simultaneous Detection and Segmentation
Cited by in corpus (15)
- SoundNet: Learning Sound Representations from Unlabeled Video
- Temporal 3D ConvNets: New Architecture and Transfer Learning for Video Classification
- DSOD: Learning Deeply Supervised Object Detectors from Scratch
- See, Hear, and Read: Deep Aligned Representations
- Stacked Generative Adversarial Networks
- Data Distillation: Towards Omni-Supervised Learning
- Amalgamating Knowledge towards Comprehensive Classification
- RGB-based 3D Hand Pose Estimation via Privileged Learning with Depth Images
- Cooperative Learning with Visual Attributes
- The Role of Context Selection in Object Detection
- Cross-lingual Distillation for Text Classification
- SurfConv: Bridging 3D and 2D Convolution for RGBD Images
- Efficient Fusion of Sparse and Complementary Convolutions
- Using Cross-Model EgoSupervision to Learn Cooperative Basketball Intention
- RGB-D Salient Object Detection Based on Discriminative Cross-modal Transfer Learning