Learning Spatial Regularization with Image-level Supervisions for Multi-label Image Classification
arXiv:1702.05891
Abstract
Multi-label image classification is a fundamental but challenging task in computer vision. Great progress has been achieved by exploiting semantic relations between labels in recent years. However, conventional approaches are unable to model the underlying spatial relations between labels in multi-label images, because spatial annotations of the labels are generally not provided. In this paper, we propose a unified deep neural network that exploits both semantic and spatial relations between labels with only image-level supervisions. Given a multi-label image, our proposed Spatial Regularization Network (SRN) generates attention maps for all labels and captures the underlying relations between them via learnable convolutions. By aggregating the regularized classification results with original results by a ResNet-101 network, the classification performance can be consistently improved. The whole deep neural network is trained end-to-end with only image-level annotations, thus requires no additional efforts on image annotations. Extensive evaluations on 3 public datasets with different types of labels show that our approach significantly outperforms state-of-the-arts and has strong generalization capability. Analysis of the learned SRN model demonstrates that it can effectively capture both semantic and spatial relations of labels for improving classification performance.
References in corpus (12)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Going Deeper with Convolutions
- CNN: Single-label to Multi-label
- Multiple Object Recognition with Visual Attention
- T-CNN: Tubelets with Convolutional Neural Networks for Object Detection from Videos
- Towards Good Practices for Very Deep Two-Stream ConvNets
- Deep Convolutional Ranking for Multilabel Image Annotation
- Contextual Action Recognition with R*CNN
- Learning Structured Inference Neural Networks with Label Relations
- Exploit Bounding Box Annotations for Multi-label Object Recognition
Cited by in corpus (7)
- ChestNet: A Deep Neural Network for Classification of Thoracic Diseases on Chest Radiography
- GM-MLIC: Graph Matching based Multi-Label Image Classification
- Learning Social Image Embedding with Deep Multimodal Attention Networks
- Pedestrian Attribute Recognition in Video Surveillance Scenarios Based on View-attribute Attention Localization
- Learning Spatial Pyramid Attentive Pooling in Image Synthesis and Image-to-Image Translation
- ELASTIC: Improving CNNs with Dynamic Scaling Policies
- Adversarial Learning of Label Dependency: A Novel Framework for Multi-class Classification