2 papers
cs.CV2026
IntentQA: Intent Question Answering in Videos by Cognitive Context Reasoning
Jiapeng Li, Ping Wei, Wenjuan Han +2
Video understanding requires intelligent agents to transcend mere recognition of visual facts and comprehend the underlying intents behind human actions (often termed the "dark mat…
cs.RO2026
UniLM-Nav: A Unified Framework for Zero-Shot Last-Mile Navigation
Zhuofan Zhang, Tianxu Wang, Guoxi Zhang +6
Mobile manipulation requires a robot to navigate to a target object or receptacle and then perform intended manipulation. However, reaching the vicinity of the target does not guar…