RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing Data
arXiv:2210.12634 · doi:10.1109/TGRS.2023.3250471
Abstract
In this paper, we introduce the task of visual grounding for remote sensing data (RSVG). RSVG aims to localize the referred objects in remote sensing (RS) images with the guidance of natural language. To retrieve rich information from RS imagery using natural language, many research tasks, like RS image visual question answering, RS image captioning, and RS image-text retrieval have been investigated a lot. However, the object-level visual grounding on RS images is still under-explored. Thus, in this work, we propose to construct the dataset and explore deep learning models for the RSVG task. Specifically, our contributions can be summarized as follows. 1) We build the new large-scale benchmark dataset of RSVG, termed RSVGD, to fully advance the research of RSVG. This new dataset includes image/expression/box triplets for training and evaluating visual grounding models. 2) We benchmark extensive state-of-the-art (SOTA) natural image visual grounding methods on the constructed RSVGD dataset, and some insightful analyses are provided based on the results. 3) A novel transformer-based Multi-Level Cross-Modal feature learning (MLCM) module is proposed. Remotely-sensed images are usually with large scale variations and cluttered backgrounds. To deal with the scale-variation problem, the MLCM module takes advantage of multi-scale visual features and multi-granularity textual embeddings to learn more discriminative representations. To cope with the cluttered background problem, MLCM adaptively filters irrelevant noise and enhances salient features. In this way, our proposed model can incorporate more effective multi-level and multi-modal features to boost performance. Furthermore, this work also provides useful insights for developing better RSVG models. The dataset and code will be publicly available at https://github.com/ZhanYang-nwpu/RSVG-pytorch.
12 pages, 10 figures
References in corpus (5)
- Exploring a Fine-Grained Multiscale Method for Cross-Modal Remote Sensing Image Retrieval
- Learning to Compose and Reason with Language Tree Structures for Visual Grounding
- From Easy to Hard: Learning Language-guided Curriculum for Visual Question Answering on Remote Sensing Data
- EarthNets: Empowering AI in Earth Observation
- Deep Unsupervised Contrastive Hashing for Large-Scale Cross-Modal Text-Image Retrieval in Remote Sensing
Cited by in corpus (8)
- RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language Model for Remote Sensing
- Remote Sensing SpatioTemporal Vision-Language Models: A Comprehensive Survey
- Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives
- RSTeller: Scaling Up Visual Language Modeling in Remote Sensing with Rich Linguistic Semantics from Openly Available Data and Large Language Models
- AeroReformer: Aerial Referring Transformer for UAV-based Referring Image Segmentation
- RSRefSeg 2: Decoupling Referring Remote Sensing Image Segmentation with Foundation Models
- Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection
- TinyRS-R1: Compact Multimodal Language Model for Remote Sensing