12 papers
MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos
Arushi Goel, Sreyan Ghosh, Vatsal Agarwal +16
Multimodal Large Language Models (MLLMs) have shown strong performance in visual and audio understanding when evaluated in isolation. However, their ability to jointly reason over…
VEGA: Learning Navigation VLAs from In-the-Wild Egocentric Video with Geometric Trajectory Supervision
Gershom Seneviratne, Yohan Abeysinghe, Jianyu An +3
We introduce VEGA, an approach for training navigation VisionLanguage-Action (VLA) models from unlabeled egocentric navigation videos. Internet-scale egocentric videos provide a sc…
A Closer Look at Failure Modes in Temporal Understanding of Large Audio-Language Models
Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh +3
Large Audio Language Models (LALMs) achieve strong performance on a variety of audio understanding tasks but continue to struggle with temporal reasoning, a fundamental capability…
Act on What You See: Unlocking Safe Social Navigation in Vision-Language-Action Models
Qingzi Wang, Xiyang Wu, Guangyao Shi +3
Safe social navigation requires robots to distinguish people from ordinary obstacles and to react before danger becomes imminent. We show that pretrained Vision-Language-Action (VL…
Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks
Xiyang Wu, Zongxia Li, Guangyao Shi +5
Long horizon interactive environments are a testbed for evaluating agents skill usage abilities. These environments demand multi step reasoning, the chaining of multiple skills ove…
SABER: A Stealthy Agentic Black-Box Attack Framework for Vision-Language-Action Models
Xiyang Wu, Guangyao Shi, Qingzi Wang +3
Vision-language-action (VLA) models enable robots to follow natural-language instructions grounded in visual observations, but the instruction channel also introduces a critical vu…