243 citations · 406 across the 32 of their papers we have counts for
55 papers
UniGeo: A Multi-modal Large Language Model for Text-Guided Cross-View Geo-Localization
Jiahao Wen, Hang Yu, Zhedong Zheng
Text-guided drone geo-localization aims to identify a target region in a large-scale image gallery from a natural-language description. Existing methods mainly formulate this task…
The 10th AI City Challenge
Zheng Tang, Shuo Wang, David C. Anastasiu +34
The 10th AI City Challenge, held with ECCV 2026, marks a decade of community benchmarking for intelligent transportation, smart cities, and physical AI. Since its 2017 start with v…
Can Vision-Language Models Think from the Sky? Unifying UAV Reasoning and Generation
Jintao Sun, Gangyi Ding, Donglin Di +2
Vision-Language Models have achieved strong progress in ground-view visual understanding, yet they remain brittle in high-altitude Unmanned Aerial Vehicle scenes, where objects are…
Uncertainty-Aware Trajectory Prediction: A Unified Framework Harnessing Positional and Semantic Uncertainties
Jintao Sun, Hu Zhang, Gangyi Ding +1
Trajectory prediction seeks to forecast the future motion of dynamic entities, such as vehicles and pedestrians, given a temporal horizon of historical movement data and environmen…
VSearcher: Long-Horizon Multimodal Search Agent via Reinforcement Learning
Ruiyang Zhang, Qianguo Sun, Chao Song +2
Large models are increasingly becoming autonomous agents that interact with real-world environments and use external tools to augment their static capabilities. However, most recen…
SketchThinker-R1: Towards Efficient Sketch-Style Reasoning in Large Multimodal Models
Ruiyang Zhang, Dongzhan Zhou, Zhedong Zheng
Despite the empirical success of extensive, step-by-step reasoning in large multimodal models, long reasoning processes inevitably incur substantial computational overhead, i.e., i…