2 papers
cs.CV2026
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing
Kaiwen Jing, Ruixu Jia, Bingyao Li +3
Recent advances in Multimodal Large Language Models (MLLMs) have significantly improved remote sensing (RS) multimodal understanding. Language-conditioned segmentation is crucial f…
cs.CV2025
GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing
Ruizhe Ou, Yuan Hu, Fan Zhang +2
Multi-modal large language models (MLLMs) have achieved remarkable success in image- and region-level remote sensing (RS) image understanding tasks, such as image captioning, visua…