imitation learning 1latent action learning 1robot manipulation 1video pretraining 1vision-language models 1
From the 1 of 6 linked papers with an AI index.
Showing cs.ROShow all
2 papers · 1 filter
cs.RO2026
WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos
Jiahao Liu, Zhongpu Xia, Shuai Tian +13
WALA is a framework that learns executable latent actions for robot manipulation by pretraining on both action‑labeled demonstrations and unlabeled videos, predicting future change…
cs.RO2026
PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance
Yupeng Zheng, Xiang Li, Songen Gu +12
Recent advances in Vision-Language-Action (VLA) models have opened new avenues for robot manipulation, yet existing methods exhibit limited efficiency and a lack of high-level know…