Self-supervised remote sensing feature learning: Learning Paradigms, Challenges, and Future Works
arXiv:2211.08129 · doi:10.1109/TGRS.2023.3276853
Abstract
Deep learning has achieved great success in learning features from massive remote sensing images (RSIs). To better understand the connection between feature learning paradigms (e.g., unsupervised feature learning (USFL), supervised feature learning (SFL), and self-supervised feature learning (SSFL)), this paper analyzes and compares them from the perspective of feature learning signals, and gives a unified feature learning framework. Under this unified framework, we analyze the advantages of SSFL over the other two learning paradigms in RSIs understanding tasks and give a comprehensive review of the existing SSFL work in RS, including the pre-training dataset, self-supervised feature learning signals, and the evaluation methods. We further analyze the effect of SSFL signals and pre-training data on the learned features to provide insights for improving the RSI feature learning. Finally, we briefly discuss some open problems and possible research directions.
24 pages, 11 figures, 3 tables
References in corpus (27)
- Bootstrap your own latent: A new approach to self-supervised Learning
- Deep learning in remote sensing: a review
- Learning Representations by Maximizing Mutual Information Across Views
- Recent Advances in Domain Adaptation for the Classification of Remote Sensing Data
- Decomposing Motion and Content for Natural Video Sequence Prediction
- Masked Autoencoders As Spatiotemporal Learners
- An Empirical Study of Remote Sensing Pretraining
- BigEarthNet-MM: A Large Scale Multi-Modal Multi-Label Benchmark Archive for Remote Sensing Image Classification and Retrieval
- Remote Sensing Cross-Modal Text-Image Retrieval Based on Global and Local Information
- A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
- Self-supervised Pretraining of Visual Features in the Wild
- Self-labelling via simultaneous clustering and representation learning
- Augmentation-Free Graph Contrastive Learning of Invariant-Discriminative Representations
- ConvMAE: Masked Convolution Meets Masked Autoencoders
- CliqueCNN: Deep Unsupervised Exemplar Learning
- Self-Supervised Learning for Invariant Representations from Multi-Spectral and SAR Images
- False: False Negative Samples Aware Contrastive Learning for Semantic Segmentation of High-Resolution Remote Sensing Image
- Contrastive Multiview Coding with Electro-optics for SAR Semantic Segmentation
- SSL4EO-S12: A Large-Scale Multi-Modal, Multi-Temporal Dataset for Self-Supervised Learning in Earth Observation
- Deep Unsupervised Contrastive Hashing for Large-Scale Cross-Modal Text-Image Retrieval in Remote Sensing
- Mugs: A Multi-Granular Self-Supervised Learning Framework
- TOV: The Original Vision Model for Optical Remote Sensing Image Understanding via Self-supervised Learning
- Aerial Scene Parsing: From Tile-level Scene Classification to Pixel-wise Semantic Labeling
- Self-Supervision, Remote Sensing and Abstraction: Representation Learning Across 3 Million Locations
- Self-supervised Hyperspectral Image Restoration using Separable Image Prior
- SeasoNet: A Seasonal Scene Classification, segmentation and Retrieval dataset for satellite Imagery over Germany
- Learning crop type mapping from regional label proportions in large-scale SAR and optical imagery
Cited by in corpus (4)
- GraSS: Contrastive Learning with Gradient Guided Sampling Strategy for Remote Sensing Image Semantic Segmentation
- Fus-MAE: A cross-attention-based data fusion approach for Masked Autoencoders in remote sensing
- Learning with less: label-efficient land cover classification at very high spatial resolution using self-supervised deep learning
- Multi-encoder ConvNeXt Network with Smooth Attentional Feature Fusion for Multispectral Semantic Segmentation