Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives
arXiv:2505.14361 · doi:10.1109/MGRS.2025.3572702
Abstract
Vision-language modeling (VLM) aims to bridge the information gap between images and natural language. Under the new paradigm of first pre-training on massive image-text pairs and then fine-tuning on task-specific data, VLM in the remote sensing domain has made significant progress. The resulting models benefit from the absorption of extensive general knowledge and demonstrate strong performance across a variety of remote sensing data analysis tasks. Moreover, they are capable of interacting with users in a conversational manner. In this paper, we aim to provide the remote sensing community with a timely and comprehensive review of the developments in VLM using the two-stage paradigm. Specifically, we first cover a taxonomy of VLM in remote sensing: contrastive learning, visual instruction tuning, and text-conditioned image generation. For each category, we detail the commonly used network architecture and pre-training objectives. Second, we conduct a thorough review of existing works, examining foundation models and task-specific adaptation methods in contrastive-based VLM, architectural upgrades, training strategies and model capabilities in instruction-based VLM, as well as generative foundation models with their representative downstream applications. Third, we summarize datasets used for VLM pre-training, fine-tuning, and evaluation, with an analysis of their construction methodologies (including image sources and caption generation) and key properties, such as scale and task adaptability. Finally, we conclude this survey with insights and discussions on future research directions: cross-modal representation alignment, vague requirement comprehension, explanation-driven model reliability, continually scalable model capabilities, and large-scale datasets featuring richer modalities and greater challenges.
Accepted by IEEE Geoscience and Remote Sensing Magazine
References in corpus (32)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Learning to Prompt for Vision-Language Models
- Remote Sensing Image Scene Classification: Benchmark and State of the Art
- Object Detection in Optical Remote Sensing Images: A Survey and A New Benchmark
- AID: A Benchmark Dataset for Performance Evaluation of Aerial Scene Classification
- ESRGAN: Enhanced Super-Resolution Generative Adversarial Networks
- DeepGlobe 2018: A Challenge to Parse the Earth through Satellite Images
- Remote Sensing Image Change Detection with Transformers
- Exploring Models and Data for Remote Sensing Image Caption Generation
- PatternNet: A Benchmark Dataset for Performance Evaluation of Remote Sensing Image Retrieval
- BigEarthNet: A Large-Scale Benchmark Archive For Remote Sensing Image Understanding
- RSVQA: Visual Question Answering for Remote Sensing Data
- HIT-UAV: A high-altitude infrared thermal dataset for Unmanned Aerial Vehicle-based object detection
- S2Looking: A Satellite Side-Looking Dataset for Building Change Detection
- Exploring a Fine-Grained Multiscale Method for Cross-Modal Remote Sensing Image Retrieval
- RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing Data
- Enabling Country-Scale Land Cover Mapping with Meter-Resolution Satellite Imagery
- RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language Model for Remote Sensing
- Building Damage Annotation on Post-Hurricane Satellite Imagery Based on Convolutional Neural Networks
- SpectralDiff: A Generative Framework for Hyperspectral Image Classification with Diffusion Models
- Learning to Rank Question Answer Pairs with Holographic Dual LSTM Architecture
- RSDiff: Remote Sensing Image Generation from Text Using Diffusion Model
- Diffusion Models Meet Remote Sensing: Principles, Methods, and Perspectives
- Synthesizing Optical and SAR Imagery From Land Cover Maps and Auxiliary Raster Data
- Deep Semantic-Visual Alignment for Zero-Shot Remote Sensing Image Scene Classification
- Text2Seg: Remote Sensing Image Semantic Segmentation via Text-Guided Visual Foundation Models
- Segment Change Model (SCM) for Unsupervised Change detection in VHR Remote Sensing Images: a Case Study of Buildings
- RSTeller: Scaling Up Visual Language Modeling in Remote Sensing with Rich Linguistic Semantics from Openly Available Data and Large Language Models
- Integration of the 3D Environment for UAV Onboard Visual Object Tracking
- Exploring the Capability of Text-to-Image Diffusion Models with Structural Edge Guidance for Multi-Spectral Satellite Image Inpainting
- QuakeSet: A Dataset and Low-Resource Models to Monitor Earthquakes through Sentinel-1
- FPCD: An Open Aerial VHR Dataset for Farm Pond Change Detection