Self-supervised remote sensing feature learning: Learning Paradigms, Challenges, and Future Works
arXiv:2211.08129 · doi:10.1109/TGRS.2023.3276853
Abstract
Deep learning has achieved great success in learning features from massive remote sensing images (RSIs). To better understand the connection between feature learning paradigms (e.g., unsupervised feature learning (USFL), supervised feature learning (SFL), and self-supervised feature learning (SSFL)), this paper analyzes and compares them from the perspective of feature learning signals, and gives a unified feature learning framework. Under this unified framework, we analyze the advantages of SSFL over the other two learning paradigms in RSIs understanding tasks and give a comprehensive review of the existing SSFL work in RS, including the pre-training dataset, self-supervised feature learning signals, and the evaluation methods. We further analyze the effect of SSFL signals and pre-training data on the learned features to provide insights for improving the RSI feature learning. Finally, we briefly discuss some open problems and possible research directions.
24 pages, 11 figures, 3 tables
References in corpus (46)
- Bootstrap your own latent: A new approach to self-supervised Learning
- Deep learning in remote sensing: a review
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
- Domain Generalization: A Survey
- Barlow Twins: Self-Supervised Learning via Redundancy Reduction
- Learning Representations by Maximizing Mutual Information Across Views
- Object Detection in Aerial Images: A Large-Scale Benchmark and Challenges
- Recent Advances in Domain Adaptation for the Classification of Remote Sensing Data
- Self-Supervised Representation Learning: Introduction, Advances and Challenges
- Decomposing Motion and Content for Natural Video Sequence Prediction
- Florence: A New Foundation Model for Computer Vision
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
- Contrastive Learning of Medical Visual Representations from Paired Images and Text
- Hard Negative Mixing for Contrastive Learning
- Masked Autoencoders As Spatiotemporal Learners
- Towards Out-Of-Distribution Generalization: A Survey
- An Empirical Study of Remote Sensing Pretraining
- BigEarthNet-MM: A Large Scale Multi-Modal Multi-Label Benchmark Archive for Remote Sensing Image Classification and Retrieval
- KST-GCN: A Knowledge-Driven Spatial-Temporal Graph Convolutional Network for Traffic Forecasting
- Global and Local Contrastive Self-Supervised Learning for Semantic Segmentation of HR Remote Sensing Images
- Remote Sensing Cross-Modal Text-Image Retrieval Based on Global and Local Information
- A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
- Self-supervised Pretraining of Visual Features in the Wild
- Self-Supervised Multisensor Change Detection
- Self-labelling via simultaneous clustering and representation learning
- Augmentation-Free Graph Contrastive Learning of Invariant-Discriminative Representations
- Self-supervised Learning is More Robust to Dataset Imbalance
- Unsupervised Feature Learning by Autoencoder and Prototypical Contrastive Learning for Hyperspectral Classification
- ConvMAE: Masked Convolution Meets Masked Autoencoders
- CliqueCNN: Deep Unsupervised Exemplar Learning
- Self-Supervised Learning for Invariant Representations from Multi-Spectral and SAR Images
- False: False Negative Samples Aware Contrastive Learning for Semantic Segmentation of High-Resolution Remote Sensing Image
- SSL4EO-S12: A Large-Scale Multi-Modal, Multi-Temporal Dataset for Self-Supervised Learning in Earth Observation
- Contrastive Multiview Coding with Electro-optics for SAR Semantic Segmentation
- Semi-supervised learning for joint SAR and multispectral land cover classification
- Deep Unsupervised Contrastive Hashing for Large-Scale Cross-Modal Text-Image Retrieval in Remote Sensing
- Towards the Generalization of Contrastive Self-Supervised Learning
- Mugs: A Multi-Granular Self-Supervised Learning Framework
- Adversarial Masking for Self-Supervised Learning
- Self-supervised Remote Sensing Images Change Detection at Pixel-level
- Aerial Scene Parsing: From Tile-level Scene Classification to Pixel-wise Semantic Labeling
- TOV: The Original Vision Model for Optical Remote Sensing Image Understanding via Self-supervised Learning
- Self-Supervision, Remote Sensing and Abstraction: Representation Learning Across 3 Million Locations
- Self-supervised Hyperspectral Image Restoration using Separable Image Prior
- SeasoNet: A Seasonal Scene Classification, segmentation and Retrieval dataset for satellite Imagery over Germany
- Learning crop type mapping from regional label proportions in large-scale SAR and optical imagery
Cited by in corpus (4)
- GraSS: Contrastive Learning with Gradient Guided Sampling Strategy for Remote Sensing Image Semantic Segmentation
- Fus-MAE: A cross-attention-based data fusion approach for Masked Autoencoders in remote sensing
- Learning with less: label-efficient land cover classification at very high spatial resolution using self-supervised deep learning
- Multi-encoder ConvNeXt Network with Smooth Attentional Feature Fusion for Multispectral Semantic Segmentation