4 papers
Robot Trajectron V2: A Probabilistic Shared Control Framework for Navigation
Pinhao Song, Yurui Du, Ophelie Saussus +4
We propose a probabilistic shared-control solution for navigation, called Robot Trajectron V2 (RT-V2), that enables accurate intent prediction and safe, effective assistance in hum…
See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
Pengteng Li, Pinhao Song, Wuyang Li +5
We introduce SEE&TREK, the first training-free prompting framework tailored to enhance the spatial understanding of Multimodal Large Language Models (MLLMS) under vision-only const…
Mini Diffuser: Fast Multi-task Diffusion Policy Training Using Two-level Mini-batches
Yutong Hu, Pinhao Song, Kehan Wen +1
We present a method that reduces, by an order of magnitude, the time and memory needed to train multi-task vision-language robotic diffusion policies. This improvement arises from…
EventVL: Understand Event Streams via Multimodal Large Language Model
Pengteng Li, Yunfan Lu, Pinghao Song +3
The event-based Vision-Language Model (VLM) recently has made good progress for practical vision tasks. However, most of these works just utilize CLIP for focusing on traditional p…