2 papers
cs.RO2026
GDLAM: Group-Disentangled Latent Action Model for Highly Disentangled Embodied Pretraining
Jiarui Yang, Jiawei Li, Jiale Zhang +5
Latent action models (LAMs) learn action-related representations from action-free videos via self-supervised future prediction, offering a scalable paradigm for embodied intelligen…
cs.RO2026
In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use
Jiarui Yang, Wen Huang, Jiale Zhang +2
Vision-Language-Action (VLA) models have become the dominant recipe for generalist manipulation, yet they are almost universally trained by behavior cloning: a policy imitates expe…