9 papers
SnapGuard: Lightweight Prompt Injection Detection for Screenshot-Based Web Agents
Mengyao Du, Han Fang, Haokai Ma +4
Web agents have emerged as an effective paradigm for automating interactions with complex web environments, yet remain vulnerable to prompt injection attacks that embed malicious i…
GeoNav: Empowering MLLMs with dual-scale geospatial reasoning for language-goal aerial navigation
Haotian Xu, Yue Hu, Chen Gao +4
Language-goal aerial navigation requires UAVs to localize targets in the complex outdoors, such as urban blocks based on textual instructions. The indoor methods are often hard to…
CityCube: Benchmarking Cross-view Spatial Reasoning on Vision-Language Models in Urban Environments
Haotian Xu, Yue Hu, Zhengqiu Zhu +7
Cross-view spatial reasoning is essential for embodied AI, underpinning spatial understanding, mental simulation and planning in complex environments. Existing benchmarks primarily…
Sparse-Tuning: Adapting Vision Transformers with Efficient Fine-tuning and Inference
Ting Liu, Xuyang Liu, Liangtao Shi +6
Parameter-efficient fine-tuning (PEFT) has emerged as a popular solution for adapting pre-trained Vision Transformer (ViT) models to downstream applications by updating only a smal…
MaPPER: Multimodal Prior-guided Parameter Efficient Tuning for Referring Expression Comprehension
Ting Liu, Zunnan Xu, Yue Hu +3
Referring Expression Comprehension (REC), which aims to ground a local visual region via natural language, is a task that heavily relies on multimodal alignment. Most existing meth…
Towards Autonomous UAV Visual Object Search in City Space: Benchmark and Agentic Methodology
Yatai Ji, Zhengqiu Zhu, Yong Zhao +7
Aerial Visual Object Search (AVOS) tasks in urban environments require Unmanned Aerial Vehicles (UAVs) to autonomously search for and identify target objects using visual and textu…