4 papers
RAMEN: Resolution-Adjustable Multimodal Encoder for Earth Observation
Nicolas Houdré, Diego Marcos, Hugo Riffaud de Turckheim +4
Earth observation (EO) data spans a wide range of spatial, spectral, and temporal resolutions, from high-resolution optical imagery to low resolution multispectral products or rada…
Visual Question Answering on Multiple Remote Sensing Image Modalities
Hichem Boussaid, Lucrezia Tosato, Flora Weissgerber +3
The extraction of visual features is an essential step in Visual Question Answering (VQA). Building a good visual representation of the analyzed scene is indeed one of the essentia…
An LLM Agent for Automatic Geospatial Data Analysis
Yuxing Chen, Weijie Wang, Sylvain Lobry +1
Large language models (LLMs) are being used in data science code generation tasks, but they often struggle with complex sequential tasks, leading to logical errors. Their applicati…
Segmentation-guided Attention for Visual Question Answering from Remote Sensing Images
Lucrezia Tosato, Hichem Boussaid, Flora Weissgerber +3
Visual Question Answering for Remote Sensing (RSVQA) is a task that aims at answering natural language questions about the content of a remote sensing image. The visual features ex…