8 papers
More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe
Stefan Maria Ailuro, Mario Markov, Mohammad Mahdi +2
Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks. Most recent progress in this are…
From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation
Mohammad Mahdi, Nedko Savov, Danda Pani Paudel +1
Exo-to-Ego video generation aims to synthesize a first-person video from a synchronized third-person view and corresponding camera poses. While paired supervision is available, syn…
B-GRTO: Bootstrapped Group Relative Tool Optimization for Referring Segmentation
Mario Markov, Stefan Maria Ailuro, Mohammad Mahdi +2
Segmentation is a fundamental task in computer vision, underpinning pixel-level scene understanding and serving as a cornerstone for applications ranging from autonomous perception…
Accelerating Vision Foundation Models with Drop-in Depthwise Convolution
Carmelo Scribano, Mohammad Mahdi, Nedyalko Prisadnikov +5
Pretrained vision foundation models deliver strong performance across tasks with limited fine-tuning. However, their Vision Transformer (ViT) backbones impose high inference costs,…
V-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence
Jiancheng Pan, Runze Wang, Tianwen Qian +7
Cross-view object correspondence, exemplified by the representative task of ego-exo object correspondence, aims to establish consistent associations of the same object across diffe…
OSMDA: OpenStreetMap-based Domain Adaptation for Remote Sensing VLMs
Stefan Maria Ailuro, Mario Markov, Mohammad Mahdi +3
Vision-Language Models (VLMs) adapted to remote sensing rely heavily on domain-specific image-text supervision, yet high-quality annotations for satellite and aerial imagery remain…