An Empirical Study of Remote Sensing Pretraining
arXiv:2204.02825 · doi:10.1109/TGRS.2022.3176603
Abstract
Deep learning has largely reshaped remote sensing (RS) research for aerial image understanding and made a great success. Nevertheless, most of the existing deep models are initialized with the ImageNet pretrained weights. Since natural images inevitably present a large domain gap relative to aerial images, probably limiting the finetuning performance on downstream aerial scene tasks. This issue motivates us to conduct an empirical study of remote sensing pretraining (RSP) on aerial images. To this end, we train different networks from scratch with the help of the largest RS scene recognition dataset up to now -- MillionAID, to obtain a series of RS pretrained backbones, including both convolutional neural networks (CNN) and vision transformers such as Swin and ViTAE, which have shown promising performance on computer vision tasks. Then, we investigate the impact of RSP on representative downstream tasks including scene recognition, semantic segmentation, object detection, and change detection using these CNN and vision transformer backbones. Empirical study shows that RSP can help deliver distinctive performances in scene recognition tasks and in perceiving RS related semantics such as "Bridge" and "Airplane". We also find that, although RSP mitigates the data discrepancies of traditional ImageNet pretraining on RS images, it may still suffer from task discrepancies, where downstream tasks require different representations from scene recognition tasks. These findings call for further research efforts on both large-scale pretraining datasets and effective pretraining methods. The codes and pretrained models will be released at https://github.com/ViTAE-Transformer/ViTAE-Transformer-Remote-Sensing.
Accepted by IEEE TGRS, codes and pretrained models are moved to https://github.com/ViTAE-Transformer/RSP
References in corpus (17)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Rethinking Atrous Convolution for Semantic Image Segmentation
- Remote Sensing Image Change Detection with Transformers
- Remote Sensing Image Scene Classification Meets Deep Learning: Challenges, Methods, Benchmarks, and Opportunities
- High-Resolution Representations for Labeling Pixels and Regions
- R2CNN: Rotational Region CNN for Orientation Robust Scene Text Detection
- Super-resolution-based Change Detection Network with Stacked Attention Module for Images with Different Resolutions
- Stagewise Unsupervised Domain Adaptation with Adversarial Self-Training for Road Segmentation of Remote Sensing Images
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive Bias
- ResT: An Efficient Transformer for Visual Recognition
- Multi-Granularity Canonical Appearance Pooling for Remote Sensing Scene Classification
- Exploring Sequence Feature Alignment for Domain Adaptive Detection Transformers
- Arbitrary-Oriented Ship Detection through Center-Head Point Extraction
- DSP: Dual Soft-Paste for Unsupervised Domain Adaptive Semantic Segmentation
- Embedded Self-Distillation in Compact Multi-Branch Ensemble Network for Remote Sensing Scene Classification
- Invariant Deep Compressible Covariance Pooling for Aerial Scene Categorization
- RegionCL: Can Simple Region Swapping Contribute to Contrastive Learning?
Cited by in corpus (14)
- HANet: A Hierarchical Attention Network for Change Detection With Bitemporal Very-High-Resolution Remote Sensing Images
- Change Guiding Network: Incorporating Change Prior to Guide Change Detection in Remote Sensing Imagery
- RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language Model for Remote Sensing
- Current Trends in Deep Learning for Earth Observation: An Open-source Benchmark Arena for Image Classification
- CMID: A Unified Self-Supervised Learning Framework for Remote Sensing Image Understanding
- Self-supervised remote sensing feature learning: Learning Paradigms, Challenges, and Future Works
- DCN-T: Dual Context Network with Transformer for Hyperspectral Image Classification
- C2F-SemiCD: A Coarse-to-Fine Semi-Supervised Change Detection Method Based on Consistency Regularization in High-Resolution Remote Sensing Images
- Semantic-aware Dense Representation Learning for Remote Sensing Image Change Detection
- A Data-Driven Review of Remote Sensing-Based Data Fusion in Precision Agriculture from Foundational to Transformer-Based Techniques
- Single-Temporal Supervised Learning for Universal Remote Sensing Change Detection
- Be the Change You Want to See: Revisiting Remote Sensing Change Detection Practices
- BD-MSA: Body decouple VHR Remote Sensing Image Change Detection method guided by multi-scale feature information aggregation
- Ice-FMBench: A Foundation Model Benchmark for Sea Ice Type Segmentation