2 papers
cs.CV2026
MAPS: Multi-Anchor Projection Similarity for Joint Vision-Language Geo-Localization
Yutong Hu, Siyuan Tan, Shaocheng Yan +3
Humans localize places by integrating perceptual cues from vision with semantic reasoning from language, forming a scene understanding that is both intuitive and structured. Althou…
cs.CV2026
Global Cross-Modal Geo-Localization: A Million-Scale Dataset and a Physical Consistency Learning Framework
Yutong Hu, Jinhui Chen, Chaoqiang Xu +6
Cross-modal Geo-localization (CMGL) matches ground-level text descriptions with geo-tagged aerial imagery, which is crucial for pedestrian navigation and emergency response. Howeve…