Remote Sensing SpatioTemporal Vision-Language Models: A Comprehensive Survey
arXiv:2412.02573 · doi:10.1109/MGRS.2025.3598283
Abstract
The interpretation of multi-temporal remote sensing imagery is critical for monitoring Earth's dynamic processes-yet previous change detection methods, which produce binary or semantic masks, fall short of providing human-readable insights into changes. Recent advances in Vision-Language Models (VLMs) have opened a new frontier by fusing visual and linguistic modalities, enabling spatio-temporal vision-language understanding: models that not only capture spatial and temporal dependencies to recognize changes but also provide a richer interactive semantic analysis of temporal images (e.g., generate descriptive captions and answer natural-language queries). In this survey, we present the first comprehensive review of RS-STVLMs. The survey covers the evolution of models from early task-specific models to recent general foundation models that leverage powerful large language models. We discuss progress in representative tasks, such as change captioning, change question answering, and change grounding. Moreover, we systematically dissect the fundamental components and key technologies underlying these models, and review the datasets and evaluation metrics that have driven the field. By synthesizing task-level insights with a deep dive into shared architectural patterns, we aim to illuminate current achievements and chart promising directions for future research in spatio-temporal vision-language understanding for remote sensing. We will keep tracing related works at https://github.com/Chen-Yang-Liu/Awesome-RS-SpatioTemporal-VLMs
Published in IEEE Geoscience and Remote Sensing Magazine
References in corpus (31)
- Knowledge Distillation: A Survey
- Deep learning in remote sensing: a review
- Remote Sensing Image Scene Classification: Benchmark and State of the Art
- Object Detection in Optical Remote Sensing Images: A Survey and A New Benchmark
- A Survey on Large Language Model based Autonomous Agents
- Remote Sensing Image Change Detection with Transformers
- Remote Sensing Image Scene Classification Meets Deep Learning: Challenges, Methods, Benchmarks, and Opportunities
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- SpectralGPT: Spectral Remote Sensing Foundation Model
- ChangeMamba: Remote Sensing Change Detection With Spatiotemporal State Space Model
- RSVQA: Visual Question Answering for Remote Sensing Data
- HANet: A Hierarchical Attention Network for Change Detection With Bitemporal Very-High-Resolution Remote Sensing Images
- S2Looking: A Satellite Side-Looking Dataset for Building Change Detection
- From W-Net to CDGAN: Bi-temporal Change Detection via Deep Learning Techniques
- RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing Data
- Joint Spatio-Temporal Modeling for the Semantic Change Detection in Remote Sensing Images
- RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language Model for Remote Sensing
- Changes to Captions: An Attentive Network for Remote Sensing Change Captioning
- Change-Agent: Towards Interactive Comprehensive Remote Sensing Change Interpretation and Analysis
- Artificial intelligence to advance Earth observation: : A review of models, recent trends, and pathways forward
- Change Detection Meets Visual Question Answering
- RDP-Net: Region Detail Preserving Network for Change Detection
- MetaBalance: Improving Multi-Task Recommendations via Adapting Gradient Magnitudes of Auxiliary Tasks
- A Lightweight Sparse Focus Transformer for Remote Sensing Image Change Captioning
- TSAR-MVS: Textureless-aware Segmentation and Correlative Refinement Guided Multi-View Stereo
- Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives
- Mask Approximation Net: A Novel Diffusion Model Approach for Remote Sensing Change Captioning
- Implicit Ray-Transformers for Multi-view Remote Sensing Image Segmentation
- Robust Change Captioning in Remote Sensing: SECOND-CC Dataset and MModalCC Framework
- QuakeSet: A Dataset and Low-Resource Models to Monitor Earthquakes through Sentinel-1
- ImageRAG: Enhancing Ultra High Resolution Remote Sensing Imagery Analysis with ImageRAG