HD-CNN: Hierarchical Deep Convolutional Neural Network for Large Scale Visual Recognition
arXiv:1410.0736
Abstract
In image classification, visual separability between different object categories is highly uneven, and some categories are more difficult to distinguish than others. Such difficult categories demand more dedicated classifiers. However, existing deep convolutional neural networks (CNN) are trained as flat N-way classifiers, and few efforts have been made to leverage the hierarchical structure of categories. In this paper, we introduce hierarchical deep CNNs (HD-CNNs) by embedding deep CNNs into a category hierarchy. An HD-CNN separates easy classes using a coarse category classifier while distinguishing difficult classes using fine category classifiers. During HD-CNN training, component-wise pretraining is followed by global finetuning with a multinomial logistic loss regularized by a coarse category consistency term. In addition, conditional executions of fine category classifiers and layer parameter compression make HD-CNNs scalable for large-scale visual recognition. We achieve state-of-the-art results on both CIFAR100 and large-scale ImageNet 1000-class benchmark datasets. In our experiments, we build up three different HD-CNNs and they lower the top-1 error of the standard CNNs by 2.65%, 3.1% and 1.1%, respectively.
Add new results on ImageNet using VGG-16-layer building block net
References in corpus (10)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- DeepPose: Human Pose Estimation via Deep Neural Networks
- Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition
- Going Deeper with Convolutions
- Deeply-Supervised Nets
- CNN: Single-label to Multi-label
- Stochastic Pooling for Regularization of Deep Convolutional Neural Networks
- Deep Convolutional Ranking for Multilabel Image Annotation
- Deep Networks with Internal Selective Attention through Feedback Connections
- Improving Deep Neural Networks with Probabilistic Maxout Units
Cited by in corpus (12)
- Beyond One-hot Encoding: lower dimensional target embedding
- Recent Advances in Convolutional Neural Networks
- Naive-Deep Face Recognition: Touching the Limit of LFW Benchmark or Not?
- Deep Contrast Learning for Salient Object Detection
- Pose-Invariant Face Alignment with a Single CNN
- Knowledge Concentration: Learning 100K Object Classifiers in a Single CNN
- Borrowing Treasures from the Wealthy: Deep Transfer Learning through Selective Joint Fine-tuning
- The iNaturalist Species Classification and Detection Dataset
- MOS: Towards Scaling Out-of-distribution Detection for Large Semantic Space
- SketchParse : Towards Rich Descriptions for Poorly Drawn Sketches using Multi-Task Hierarchical Deep Networks
- Basic Level Categorization Facilitates Visual Object Recognition
- HUSE: Hierarchical Universal Semantic Embeddings