Exploring Masked Autoencoders for Sensor-Agnostic Image Retrieval in Remote Sensing
arXiv:2401.07782 · doi:10.1109/TGRS.2024.3517150
Abstract
Self-supervised learning through masked autoencoders (MAEs) has recently attracted great attention for remote sensing (RS) image representation learning, and thus embodies a significant potential for content-based image retrieval (CBIR) from ever-growing RS image archives. However, the existing MAE based CBIR studies in RS assume that the considered RS images are acquired by a single image sensor, and thus are only suitable for uni-modal CBIR problems. The effectiveness of MAEs for cross-sensor CBIR, which aims to search semantically similar images across different image modalities, has not been explored yet. In this paper, we take the first step to explore the effectiveness of MAEs for sensor-agnostic CBIR in RS. To this end, we present a systematic overview on the possible adaptations of the vanilla MAE to exploit masked image modeling on multi-sensor RS image archives (denoted as cross-sensor masked autoencoders [CSMAEs]) in the context of CBIR. Based on different adjustments applied to the vanilla MAE, we introduce different CSMAE models. We also provide an extensive experimental analysis of these CSMAE models. We finally derive a guideline to exploit masked image modeling for uni-modal and cross-modal CBIR problems in RS. The code of this work is publicly available at https://github.com/jakhac/CSMAE.
Accepted at the IEEE Transactions on Geoscience and Remote Sensing. Our code is available at https://github.com/jakhac/CSMAE
References in corpus (14)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- YFCC100M: The New Data in Multimedia Research
- Learning Representations by Maximizing Mutual Information Across Views
- BigEarthNet-MM: A Large Scale Multi-Modal Multi-Label Benchmark Archive for Remote Sensing Image Classification and Retrieval
- SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite Imagery
- CMID: A Unified Self-Supervised Learning Framework for Remote Sensing Image Understanding
- A Billion-scale Foundation Model for Remote Sensing Images
- CMIR-NET : A Deep Learning Based Model For Cross-Modal Retrieval In Remote Sensing
- Exploring Plain Vision Transformer Backbones for Object Detection
- Masked Vision and Language Modeling for Multi-modal Representation Learning
- Neural Plasticity-Inspired Multimodal Foundation Model for Earth Observation
- A Novel Self-Supervised Cross-Modal Image Retrieval Method In Remote Sensing
- Generative Reasoning Integrated Label Noise Robust Deep Image Representation Learning
- USat: A Unified Self-Supervised Encoder for Multi-Sensor Satellite Imagery