Semantic-aware Dense Representation Learning for Remote Sensing Image Change Detection
arXiv:2205.13769 · doi:10.1109/TGRS.2022.3203769
Abstract
Supervised deep learning models depend on massive labeled data. Unfortunately, it is time-consuming and labor-intensive to collect and annotate bitemporal samples containing desired changes. Transfer learning from pre-trained models is effective to alleviate label insufficiency in remote sensing (RS) change detection (CD). We explore the use of semantic information during pre-training. Different from traditional supervised pre-training that learns the mapping from image to label, we incorporate semantic supervision into the self-supervised learning (SSL) framework. Typically, multiple objects of interest (e.g., buildings) are distributed in various locations in an uncurated RS image. Instead of manipulating image-level representations via global pooling, we introduce point-level supervision on per-pixel embeddings to learn spatially-sensitive features, thus benefiting downstream dense CD. To achieve this, we obtain multiple points via class-balanced sampling on the overlapped area between views using the semantic mask. We learn an embedding space where background and foreground points are pushed apart, and spatially aligned points across views are pulled together. Our intuition is the resulting semantically discriminative representations invariant to irrelevant changes (illumination and unconcerned land covers) may help change recognition. We collect large-scale image-mask pairs freely available in the RS community for pre-training. Extensive experiments on three CD datasets verify the effectiveness of our method. Ours significantly outperforms ImageNet pre-training, in-domain supervision, and several SSL methods. Empirical results indicate our pre-training improves the generalization and data efficiency of the CD model. Notably, we achieve competitive results using 20% training data than baseline (random initialization) using 100% data. Our code is available.
18 pages, 7 figures. Accepted article by IEEE TGRS
References in corpus (12)
- Bootstrap your own latent: A new approach to self-supervised Learning
- Deep learning in remote sensing: a review
- More Diverse Means Better: Multimodal Deep Learning Meets Remote Sensing Imagery Classification
- An Empirical Study of Remote Sensing Pretraining
- Building Damage Detection in Satellite Imagery Using Convolutional Neural Networks
- Revisiting Consistency Regularization for Semi-supervised Change Detection in Remote Sensing Images
- Self-Supervised Learning for Invariant Representations from Multi-Spectral and SAR Images
- Contrastive Multiview Coding with Electro-optics for SAR Semantic Segmentation
- Deep Active Learning in Remote Sensing for data efficient Change Detection
- TOV: The Original Vision Model for Optical Remote Sensing Image Understanding via Self-supervised Learning
- Embedding Earth: Self-supervised contrastive pre-training for dense land cover classification
- Active learning for interactive satellite image change detection
Cited by in corpus (5)
- HANet: A Hierarchical Attention Network for Change Detection With Bitemporal Very-High-Resolution Remote Sensing Images
- Continuous Remote Sensing Image Super-Resolution based on Context Interaction in Implicit Function Space
- Continuous Cross-resolution Remote Sensing Image Change Detection
- EfficientCD: A New Strategy For Change Detection Based With Bi-temporal Layers Exchanged
- AdaSemiCD: An Adaptive Semi-Supervised Change Detection Method Based on Pseudo-Label Evaluation