4 papers
BGG: Bridging the Geometric Gap between Cross-View images by Vision Foundation Model Adaptation for Geo-Localization
Wei Wang, Dou Quan, Ning Huyan +4
Geometric differences between cross-view images, such as drone and satellite views, significantly increase the challenge of Cross-View Geo-Localization (CVGL), which aims to acquir…
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
Shida Gao, Feng Xue, Xiangfeng Wang +8
Multimodal large language models (MLLMs) are rapidly expanding from general video understanding to finer-grained understanding such as spatio-temporal video grounding (STVG) and re…
Multi-Expert Learning Framework with the State Space Model for Optical and SAR Image Registration
Wei Wang, Dou Quan, Ning Huyan +4
Optical and Synthetic Aperture Radar (SAR) image registration is crucial for multi-modal image fusion and applications. However, several challenges limit the performance of existin…
CLNet: Cross-View Correspondence Makes a Stronger Geo-Localizationer
Xianwei Cao, Dou Quan, Shuang Wang +4
Image retrieval-based cross-view geo-localization (IRCVGL) aims to match images captured from significantly different viewpoints, such as satellite and street-level images. Existin…