Improving Semantic Segmentation of Aerial Images Using Patch-based Attention
arXiv:1911.08877 · doi:10.1109/TGRS.2020.2994150
Abstract
The trade-off between feature representation power and spatial localization accuracy is crucial for the dense classification/semantic segmentation of aerial images. High-level features extracted from the late layers of a neural network are rich in semantic information, yet have blurred spatial details; low-level features extracted from the early layers of a network contain more pixel-level information, but are isolated and noisy. It is therefore difficult to bridge the gap between high and low-level features due to their difference in terms of physical information content and spatial distribution. In this work, we contribute to solve this problem by enhancing the feature representation in two ways. On the one hand, a patch attention module (PAM) is proposed to enhance the embedding of context information based on a patch-wise calculation of local attention. On the other hand, an attention embedding module (AEM) is proposed to enrich the semantic information of low-level features by embedding local focus from high-level features. Both of the proposed modules are light-weight and can be applied to process the extracted features of convolutional neural networks (CNNs). Experiments show that, by integrating the proposed modules into the baseline Fully Convolutional Network (FCN), the resulting local attention network (LANet) greatly improves the performance over the baseline and outperforms other attention based methods on two aerial image datasets.
[J]. IEEE Transactions on Geoscience and Remote Sensing, 2020
References in corpus (1)
Cited by in corpus (21)
- UNetFormer: A UNet-like Transformer for Efficient Semantic Segmentation of Remote Sensing Urban Scene Imagery
- A Review on Deep Learning in UAV Remote Sensing
- Multi-Attention-Network for Semantic Segmentation of Fine Resolution Remote Sensing Images
- A2-FPN for Semantic Segmentation of Fine-Resolution Remotely Sensed Images
- Bi-Temporal Semantic Reasoning for the Semantic Change Detection in HR Remote Sensing Images
- An Empirical Study of Remote Sensing Pretraining
- Global and Local Contrastive Self-Supervised Learning for Semantic Segmentation of HR Remote Sensing Images
- Looking Outside the Window: Wide-Context Transformer for the Semantic Segmentation of High-Resolution Remote Sensing Images
- Joint Spatio-Temporal Modeling for the Semantic Change Detection in Remote Sensing Images
- BDANet: Multiscale Convolutional Neural Network with Cross-directional Attention for Building Damage Assessment from Satellite Images
- Adversarial Shape Learning for Building Extraction in VHR Remote Sensing Images
- Txt2Img-MHN: Remote Sensing Image Generation from Text Using Modern Hopfield Networks
- MP-ResNet: Multi-path Residual Network for the Semantic segmentation of High-Resolution PolSAR Images
- MKANet: A Lightweight Network with Sobel Boundary Loss for Efficient Land-cover Classification of Satellite Remote Sensing Imagery
- Semantic Labeling of High Resolution Images Using EfficientUNets and Transformers
- MetaSegNet: Metadata-collaborative Vision-Language Representation Learning for Semantic Segmentation of Remote Sensing Images
- SmartMem: Layout Transformation Elimination and Adaptation for Efficient DNN Execution on Mobile
- LoLA-SpecViT: Local Attention SwiGLU Vision Transformer with LoRA for Hyperspectral Imaging
- Dual Attention GANs for Semantic Image Synthesis
- Audio-Visual Event Localization via Recursive Fusion by Joint Co-Attention
- IRSAMap:Towards Large-Scale, High-Resolution Land Cover Map Vectorization