Relation Network for Multi-label Aerial Image Classification
arXiv:1907.07274 · doi:10.1109/TGRS.2019.2963364
Abstract
Multi-label classification plays a momentous role in perceiving intricate contents of an aerial image and triggers several related studies over the last years. However, most of them deploy few efforts in exploiting label relations, while such dependencies are crucial for making accurate predictions. Although an LSTM layer can be introduced to modeling such label dependencies in a chain propagation manner, the efficiency might be questioned when certain labels are improperly inferred. To address this, we propose a novel aerial image multi-label classification network, attention-aware label relational reasoning network. Particularly, our network consists of three elemental modules: 1) a label-wise feature parcel learning module, 2) an attentional region extraction module, and 3) a label relational inference module. To be more specific, the label-wise feature parcel learning module is designed for extracting high-level label-specific features. The attentional region extraction module aims at localizing discriminative regions in these features and yielding attentional label-specific features. The label relational inference module finally predicts label existences using label relations reasoned from outputs of the previous module. The proposed network is characterized by its capacities of extracting discriminative label-wise features in a proposal-free way and reasoning about label relations naturally and interpretably. In our experiments, we evaluate the proposed model on the UCM multi-label dataset and a newly produced dataset, AID multi-label dataset. Quantitative and qualitative results on these two datasets demonstrate the effectiveness of our model. To facilitate progress in the multi-label aerial image classification, the AID multi-label dataset will be made publicly available.
References in corpus (4)
- Deep learning in remote sensing: a review
- Understanding urban landuse from the above and ground perspectives: a deep learning, multimodal solution
- Fine-Grained Object Recognition and Zero-Shot Learning in Remote Sensing Imagery
- A Novel Multi-Attention Driven System For Multi-Label Remote Sensing Image Classification
Cited by in corpus (12)
- Remote Sensing Image Scene Classification Meets Deep Learning: Challenges, Methods, Benchmarks, and Opportunities
- Current Trends in Deep Learning for Earth Observation: An Open-source Benchmark Arena for Image Classification
- Semantics-Consistent Representation Learning for Remote Sensing Image-Voice Retrieval
- All Grains, One Scheme (AGOS): Learning Multi-grain Instance Representation for Aerial Scene Classification
- Semi-Supervised Building Footprint Generation with Feature and Output Consistency Training
- Multi-label Image Classification using Adaptive Graph Convolutional Networks: from a Single Domain to Multiple Domains
- MUS-CDB: Mixed Uncertainty Sampling with Class Distribution Balancing for Active Annotation in Aerial Object Detection
- MultiScene: A Large-scale Dataset and Benchmark for Multi-scene Recognition in Single Aerial Images
- Rethinking Crowdsourcing Annotation: Partial Annotation with Salient Labels for Multi-Label Image Classification
- SCIDA: Self-Correction Integrated Domain Adaptation from Single- to Multi-label Aerial Images
- CG-Net: Conditional GIS-aware Network for Individual Building Segmentation in VHR SAR Images
- Adaptive Gradient Calibration for Single-Positive Multi-Label Learning in Remote Sensing Image Scene Classification