Looking Outside the Window: Wide-Context Transformer for the Semantic Segmentation of High-Resolution Remote Sensing Images
arXiv:2106.15754 · doi:10.1109/TGRS.2022.3168697
Abstract
Long-range contextual information is crucial for the semantic segmentation of High-Resolution (HR) Remote Sensing Images (RSIs). However, image cropping operations, commonly used for training neural networks, limit the perception of long-range contexts in large RSIs. To overcome this limitation, we propose a Wide-Context Network (WiCoNet) for the semantic segmentation of HR RSIs. Apart from extracting local features with a conventional CNN, the WiCoNet has an extra context branch to aggregate information from a larger image area. Moreover, we introduce a Context Transformer to embed contextual information from the context branch and selectively project it onto the local features. The Context Transformer extends the Vision Transformer, an emerging kind of neural network, to model the dual-branch semantic correlations. It overcomes the locality limitation of CNNs and enables the WiCoNet to see the bigger picture before segmenting the land-cover/land-use (LCLU) classes. Ablation studies and comparative experiments conducted on several benchmark datasets demonstrate the effectiveness of the proposed method. In addition, we present a new Beijing Land-Use (BLU) dataset. This is a large-scale HR satellite dataset with high-quality and fine-grained reference labels, which can facilitate future studies in this field.
References in corpus (10)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation
- Transformers in Vision: A Survey
- Fully Convolutional Networks for Semantic Segmentation
- Remote Sensing Image Change Detection with Transformers
- Object Detectors Emerge in Deep Scene CNNs
- Dense semantic labeling of sub-decimeter resolution images with convolutional neural networks
- Adversarial Shape Learning for Building Extraction in VHR Remote Sensing Images
- MP-ResNet: Multi-path Residual Network for the Semantic segmentation of High-Resolution PolSAR Images
- Transformer-Based Source-Free Domain Adaptation
Cited by in corpus (6)
- Joint Spatio-Temporal Modeling for the Semantic Change Detection in Remote Sensing Images
- A Survey of Sample-Efficient Deep Learning for Change Detection in Remote Sensing: Tasks, Strategies, and Challenges
- LMFNet: An Efficient Multimodal Fusion Approach for Semantic Segmentation in High-Resolution Remote Sensing
- MKANet: A Lightweight Network with Sobel Boundary Loss for Efficient Land-cover Classification of Satellite Remote Sensing Imagery
- Graph Information Bottleneck for Remote Sensing Segmentation
- SRMF: A Data Augmentation and Multimodal Fusion Approach for Long-Tail UHR Satellite Image Segmentation