7 papers
RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation
Feng Jiang, Yang Chen, Kyle Xu +8
Recent advances in large-scale video world models have enabled increasingly realistic future prediction, raising the prospect of using generated videos as scalable supervision for…
M-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering
Jiatong Ma, Longteng Guo, Yuchen Liu +4
We present M-VQA, a novel knowledge-based Visual Question Answering (VQA) benchmark, to enhance the evaluation of multimodal large language models (MLLMs) in fine-grained multi…
BiPreManip: Learning Affordance-Based Bimanual Preparatory Manipulation through Anticipatory Collaboration
Yan Shen, Feng Jiang, Zichen He +5
Many everyday objects are difficult to directly grasp (e.g., a flat iPad) or manipulate functionally (e.g., opening the cap of a pen lying on a desk). Such tasks require sequential…
Seeing is Believing (and Predicting): Context-Aware Multi-Human Behavior Prediction with Vision Language Models
Utsav Panchal, Yuchen Liu, Luigi Palmieri +2
Accurately predicting human behaviors is crucial for mobile robots operating in human-populated environments. While prior research primarily focuses on predicting actions in single…
Vision Language Models Cannot Plan, but Can They Formalize?
Muyu He, Yuxi Zheng, Yuchen Liu +7
The advancement of vision language models (VLMs) has empowered embodied agents to accomplish simple multimodal planning tasks, but not long-horizon ones requiring long sequences of…
Context-Aware Human Behavior Prediction Using Multimodal Large Language Models: Challenges and Insights
Yuchen Liu, Lino Lerch, Luigi Palmieri +4
Predicting human behavior in shared environments is crucial for safe and efficient human-robot interaction. Traditional data-driven methods to that end are pre-trained on domain-sp…