7 papers
Sentinel2Cap: A Human-Annotated Benchmark Dataset for Multimodal Remote Sensing Image Captioning
Lucrezia Tosato, Gianluca Lombardi, Ronny Hansch
Image captioning has become an important task in computer vision, enabling models to generate natural language descriptions of visual content. While several datasets exist for natu…
Checkmate: interpretable and explainable RSVQA is the endgame
Lucrezia Tosato, Christel Tartini Chappuis, Syrielle Montariol +3
Remote Sensing Visual Question Answering (RSVQA) presents unique challenges in ensuring that model decisions are both understandable and grounded in visual content. Current models…
SAR Strikes Back: A New Hope for RSVQA
Lucrezia Tosato, Flora Weissgerber, Laurent Wendling +1
Remote Sensing Visual Question Answering (RSVQA) is a task that extracts information from satellite images to answer questions in natural language, aiding image interpretation. Whi…
Visual Question Answering on Multiple Remote Sensing Image Modalities
Hichem Boussaid, Lucrezia Tosato, Flora Weissgerber +3
The extraction of visual features is an essential step in Visual Question Answering (VQA). Building a good visual representation of the analyzed scene is indeed one of the essentia…
Exploiting temporal information to detect conversational groups in videos and predict the next speaker
Lucrezia Tosato, Victor Fortier, Isabelle Bloch +1
Studies in human human interaction have introduced the concept of F formation to describe the spatial arrangement of participants during social interactions. This paper has two obj…
Can SAR improve RSVQA performance?
Lucrezia Tosato, Sylvain Lobry, Flora Weissgerber +1
Remote sensing visual question answering (RSVQA) has been involved in several research in recent years, leading to an increase in new methods. RSVQA automatically extracts informat…