activity
20242026
most citedGeoNav: Empowering MLLMs with dual-scale geospatial reasoning for language-goal aerial navigation

6 citations · 7 across the 11 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

CityCube: Benchmarking Cross-view Spatial Reasoning on Vision-Language Models in Urban Environments

Haotian Xu, Yue Hu, Zhengqiu Zhu +7

Cross-view spatial reasoning is essential for embodied AI, underpinning spatial understanding, mental simulation and planning in complex environments. Existing benchmarks primarily…

cs.CV2025

Towards Autonomous UAV Visual Object Search in City Space: Benchmark and Agentic Methodology

Yatai Ji, Zhengqiu Zhu, Yong Zhao +7

Aerial Visual Object Search (AVOS) tasks in urban environments require Unmanned Aerial Vehicles (UAVs) to autonomously search for and identify target objects using visual and textu…

cs.CV2025

SwimVG: Step-wise Multimodal Fusion and Adaption for Visual Grounding

Liangtao Shi, Ting Liu, Xiantao Hu +3

Visual grounding aims to ground an image region through natural language, which heavily relies on cross-modal alignment. Most existing methods transfer visual/linguistic knowledge…

cs.CV2024★ 1 cited

Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

Ting Liu, Liangtao Shi, Richang Hong +3

The vision tokens in multimodal large language models usually exhibit significant spatial and temporal redundancy and take up most of the input tokens, which harms their inference…

cs.CV2024

MaPPER: Multimodal Prior-guided Parameter Efficient Tuning for Referring Expression Comprehension

Ting Liu, Zunnan Xu, Yue Hu +3

Referring Expression Comprehension (REC), which aims to ground a local visual region via natural language, is a task that heavily relies on multimodal alignment. Most existing meth…

cs.CV2024

M2IST: Multi-Modal Interactive Side-Tuning for Efficient Referring Expression Comprehension

Xuyang Liu, Ting Liu, Siteng Huang +6

Referring expression comprehension (REC) is a vision-language task to locate a target object in an image based on a language expression. Fully fine-tuning general-purpose pre-train…