5 papers
AirForesight: Current-to-Future Spatial Map Imagination with Cross-Space Planning Consistency for UAV-VLN
Yutong Liu, Xiaojie Li, Mingzhu Xu +1
Unmanned Aerial Vehicle Vision-Language Navigation (UAV-VLN) requires agents to follow language instructions, infer spatial structure from sparse multi-view observations, and execu…
TimeLogic Challenge @ CVPR 2026: Strong MLLMs Meet Evidence-Seeking Agents for Temporal-Logic Video Question Answering
Zhaoyang Xu, Xusheng He, Wei Liu +2
Temporal-logic video question answering requires a model to reason about when actions occur relative to one another, such as before, after, until, since, overlap, and multi-event c…
FineBadminton: A Multi-Level Dataset for Fine-Grained Badminton Video Understanding
Xusheng He, Wei Liu, Shanshan Ma +3
Fine-grained analysis of complex and high-speed sports like badminton presents a significant challenge for Multimodal Large Language Models (MLLMs), despite their notable advanceme…
A Survey on Video Temporal Grounding with Multimodal Large Language Model
Jianlong Wu, Wei Liu, Ye Liu +4
The recent advancement in video temporal grounding (VTG) has significantly enhanced fine-grained video understanding, primarily driven by multimodal large language models (MLLMs).…
AffordGrasp: In-Context Affordance Reasoning for Open-Vocabulary Task-Oriented Grasping in Clutter
Yingbo Tang, Shuaike Zhang, Xiaoshuai Hao +4
Inferring the affordance of an object and grasping it in a task-oriented manner is crucial for robots to successfully complete manipulation tasks. Affordance indicates where and ho…