3 papers
cs.CV2025
An Efficient and Effective Encoder Model for Vision and Language Tasks in the Remote Sensing Domain
João Daniel Silva, Joao Magalhaes, Devis Tuia +1
The remote sensing community has recently seen the emergence of methods based on Large Vision and Language Models (LVLMs) that can address multiple tasks at the intersection of com…
cs.CV2025
Multilingual Training-Free Remote Sensing Image Captioning
Carlos Rebelo, Gil Rocha, João Daniel Silva +1
Remote sensing image captioning has advanced rapidly through encoder--decoder models, although the reliance on large annotated datasets and the focus on English restricts global ap…
cs.CV2024
Multilingual Vision-Language Pre-training for the Remote Sensing Domain
João Daniel Silva, Joao Magalhaes, Devis Tuia +1
Methods based on Contrastive Language-Image Pre-training (CLIP) are nowadays extensively used in support of vision-and-language tasks involving remote sensing data, such as cross-m…