4 papers
Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets?
Dingrui Wang, Zhihao Liang, Hongyuan Ye +13
While recent video world models can generate highly realistic videos, their ability to perform semantic reasoning and planning remains unclear and unquantified. We introduce Target…
Seeing is Believing (and Predicting): Context-Aware Multi-Human Behavior Prediction with Vision Language Models
Utsav Panchal, Yuchen Liu, Luigi Palmieri +2
Accurately predicting human behaviors is crucial for mobile robots operating in human-populated environments. While prior research primarily focuses on predicting actions in single…
Context-Aware Human Behavior Prediction Using Multimodal Large Language Models: Challenges and Insights
Yuchen Liu, Lino Lerch, Luigi Palmieri +4
Predicting human behavior in shared environments is crucial for safe and efficient human-robot interaction. Traditional data-driven methods to that end are pre-trained on domain-sp…
DELTA: Decomposed Efficient Long-Term Robot Task Planning using Large Language Models
Yuchen Liu, Luigi Palmieri, Sebastian Koch +2
Recent advancements in Large Language Models (LLMs) have sparked a revolution across many research fields. In robotics, the integration of common-sense knowledge from LLMs into tas…