Factors of Transferability for a Generic ConvNet Representation
arXiv:1406.5774
Abstract
Evidence is mounting that Convolutional Networks (ConvNets) are the most effective representation learning method for visual recognition tasks. In the common scenario, a ConvNet is trained on a large labeled dataset (source) and the feed-forward units activation of the trained network, at a certain layer of the network, is used as a generic representation of an input image for a task with relatively smaller training set (target). Recent studies have shown this form of representation transfer to be suitable for a wide range of target visual recognition tasks. This paper introduces and investigates several factors affecting the transferability of such representations. It includes parameters for training of the source ConvNet such as its architecture, distribution of the training data, etc. and also the parameters of feature extraction such as layer of the trained ConvNet, dimensionality reduction, etc. Then, by optimizing these factors, we show that significant improvements can be achieved on various (17) visual recognition tasks. We further show that these visual recognition tasks can be categorically ordered based on their distance from the source task such that a correlation between the performance of tasks and their distance from the source task w.r.t. the proposed factors is observed.
Extended version of the workshop paper with more experiments and updated text and title. Original CVPR15 DeepVision workshop paper title: "From Generic to Specific Deep Representations for Visual Recognition"
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Caffe: Convolutional Architecture for Fast Feature Embedding
- How transferable are features in deep neural networks?
- Going Deeper with Convolutions
- Practical recommendations for gradient-based training of deep architectures
- Analyzing the Performance of Multilayer Neural Networks for Object Recognition
- Deformable Part Models are Convolutional Neural Networks
Cited by in corpus (5)
- Convolutional Neural Networks for Medical Image Analysis: Full Training or Fine Tuning?
- Design of Efficient Deep Learning models for Determining Road Surface Condition from Roadside Camera Images and Weather Data
- The Treasure beneath Convolutional Layers: Cross-convolutional-layer Pooling for Image Classification
- Dynamic texture and scene classification by transferring deep image features
- Voronoi-based compact image descriptors: Efficient Region-of-Interest retrieval with VLAD and deep-learning-based descriptors