1 citations · 1 across the 3 of their papers we have counts for
4 papers
CityCube: Benchmarking Cross-view Spatial Reasoning on Vision-Language Models in Urban Environments
Haotian Xu, Yue Hu, Zhengqiu Zhu +7
Cross-view spatial reasoning is essential for embodied AI, underpinning spatial understanding, mental simulation and planning in complex environments. Existing benchmarks primarily…
Towards Autonomous UAV Visual Object Search in City Space: Benchmark and Agentic Methodology
Yatai Ji, Zhengqiu Zhu, Yong Zhao +7
Aerial Visual Object Search (AVOS) tasks in urban environments require Unmanned Aerial Vehicles (UAVs) to autonomously search for and identify target objects using visual and textu…
SwimVG: Step-wise Multimodal Fusion and Adaption for Visual Grounding
Liangtao Shi, Ting Liu, Xiantao Hu +3
Visual grounding aims to ground an image region through natural language, which heavily relies on cross-modal alignment. Most existing methods transfer visual/linguistic knowledge…
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Ting Liu, Liangtao Shi, Richang Hong +3
The vision tokens in multimodal large language models usually exhibit significant spatial and temporal redundancy and take up most of the input tokens, which harms their inference…