4 papers
GIFT: Guided Intermediate Feature Training via Action-Oriented Structural Supervision for Robotic Manipulation
Yupeng Zheng, Xiang Li, Songen Gu +11
Vision-language pre-training and predictive world modeling provide robot policies with rich semantic and dynamic visual features, but their native action and visual-prediction obje…
WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos
Jiahao Liu, Zhongpu Xia, Shuai Tian +13
Generalizable robot policies typically rely on action-labeled robot demonstrations, which are expensive to collect and difficult to scale. In contrast, large-scale human and robot…
4DVLT: Dynamic Scene Understanding with Worldline-Centered Vision-Language Tracking
Chaoyue Li, Boxue Yang, Shengyao Zhou +3
4D dynamic scene understanding requires grounding language to a persistent worldline that binds identity, metric 3D motion, and synchronized multi-view 2D projections. Existing par…
LMM-Track4D: Eliciting 4D Dynamic Reasoning in LMMs via Trajectory-Grounded Dialogue
Chaoyue Li, Yongxue Xu, Jie Feng +1
Recent large multimodal models (LMMs) have become increasingly capable on image and video understanding, yet still struggle to sustain 4D continuous spatiotemporal dynamic reasonin…