8 papers · 1 filter
Training-Free Open-Vocabulary Visual Grounding for Remote Sensing Images and Videos
Ke Li, Di Wang, Yongshan Zhu +5
Remote sensing visual grounding (RSVG) aims to localize a referred target in a remote sensing image or video according to a natural language expression. Existing RSVG methods usual…
OmniTraffic: A Controllable Generation Pipeline and Benchmark for Spatio-Temporal Traffic Reasoning
Maonan Wang, Zhengyan Huang, Kemou Jiang +13
Traffic scene understanding requires models to reason beyond object recognition, including lane topology, multi-view geometry, temporal evolution, and signal-phase semantics. Howev…
ProVG: Progressive Visual Grounding via Language Decoupling for Remote Sensing Imagery
Ke Li, Ting Wang, Di Wang +4
Remote sensing visual grounding (RSVG) aims to localize objects in remote sensing imagery according to natural language expressions. Previous methods typically rely on sentence-lev…
Improving Few-Shot Change Detection Visual Question Answering via Decision-Ambiguity-guided Reinforcement Fine-Tuning
Fuyu Dong, Ke Li, Di Wang +5
Change detection visual question answering (CDVQA) requires answering text queries by reasoning about semantic changes in bi-temporal remote sensing images. A straightforward appro…
A Dual-Branch Local-Global Framework for Cross-Resolution Land Cover Mapping
Peng Gao, Ke Li, Di Wang +4
Cross-resolution land cover mapping aims to produce high-resolution semantic predictions from coarse or low-resolution supervision, yet the severe resolution mismatch makes effecti…
Robust Drone-View Geo-Localization via Content-Viewpoint Disentanglement
Ke Li, Di Wang, Xiaowei Wang +4
Drone-view geo-localization (DVGL) aims to match images of the same geographic location captured from drone and satellite perspectives. Despite recent advances, DVGL remains challe…