RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language Model for Remote Sensing
arXiv:2306.11300 · doi:10.1109/TGRS.2024.3449154
Abstract
Pre-trained Vision-Language Models (VLMs) utilizing extensive image-text paired data have demonstrated unprecedented image-text association capabilities, achieving remarkable results across various downstream tasks. A critical challenge is how to make use of existing large-scale pre-trained VLMs, which are trained on common objects, to perform the domain-specific transfer for accomplishing domain-related downstream tasks. A critical challenge is how to make use of existing large-scale pre-trained VLMs, which are trained on common objects, to perform the domain-specific transfer for accomplishing domain-related downstream tasks. In this paper, we propose a new framework that includes the Domain pre-trained Vision-Language Model (DVLM), bridging the gap between the General Vision-Language Model (GVLM) and domain-specific downstream tasks. Moreover, we present an image-text paired dataset in the field of remote sensing (RS), RS5M, which has 5 million RS images with English descriptions. The dataset is obtained from filtering publicly available image-text paired datasets and captioning label-only RS datasets with pre-trained VLM. These constitute the first large-scale RS image-text paired dataset. Additionally, we fine-tuned the CLIP model and tried several Parameter-Efficient Fine-Tuning methods on RS5M to implement the DVLM. Experimental results show that our proposed dataset is highly effective for various tasks, and our model GeoRSCLIP improves upon the baseline or previous state-of-the-art model by in Zero-shot Classification (ZSC), in Remote Sensing Cross-Modal Text-Image Retrieval (RSCTIR) and in Semantic Localization (SeLo) tasks. Dataset and models have been released in: \url{https://github.com/om-ai-lab/RS5M}.
RS5M dataset v5
References in corpus (5)
- An Empirical Study of Remote Sensing Pretraining
- Exploring a Fine-Grained Multiscale Method for Cross-Modal Remote Sensing Image Retrieval
- Remote Sensing Cross-Modal Text-Image Retrieval Based on Global and Local Information
- RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing Data
- Learning to Evaluate Performance of Multi-modal Semantic Localization
Cited by in corpus (7)
- UAVs Meet LLMs: Overviews and Perspectives Toward Agentic Low-Altitude Mobility
- Remote Sensing SpatioTemporal Vision-Language Models: A Comprehensive Survey
- Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives
- Mask Approximation Net: A Novel Diffusion Model Approach for Remote Sensing Change Captioning
- RSTeller: Scaling Up Visual Language Modeling in Remote Sensing with Rich Linguistic Semantics from Openly Available Data and Large Language Models
- ImageRAG: Enhancing Ultra High Resolution Remote Sensing Imagery Analysis with ImageRAG
- Survey on Remote Sensing Scene Classification: From Traditional Methods to Large Generative AI Models