imitation learning 1latent action learning 1robot manipulation 1video pretraining 1vision-language models 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.RO2026
WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos
Jiahao Liu, Zhongpu Xia, Shuai Tian +13
WALA is a framework that learns executable latent actions for robot manipulation by pretraining on both action‑labeled demonstrations and unlabeled videos, predicting future change…
cs.CV2026
4DVLT: Dynamic Scene Understanding with Worldline-Centered Vision-Language Tracking
Chaoyue Li, Boxue Yang, Shengyao Zhou +3
4D dynamic scene understanding requires grounding language to a persistent worldline that binds identity, metric 3D motion, and synchronized multi-view 2D projections. Existing par…
cs.CV2026
LMM-Track4D: Eliciting 4D Dynamic Reasoning in LMMs via Trajectory-Grounded Dialogue
Chaoyue Li, Yongxue Xu, Jie Feng +1
Recent large multimodal models (LMMs) have become increasingly capable on image and video understanding, yet still struggle to sustain 4D continuous spatiotemporal dynamic reasonin…