2 papers
cs.RO2026
Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision
Haoyang Li, Guanlin Li, Youhe Feng +9
Cross-embodiment transfer in vision-language-action (VLA) models remains challenging because low-level state and action spaces differ fundamentally across robot platforms. We obser…
cs.RO2026
From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models
Yihan Lin, Haoyang Li, Yang Li +4
Latent actions serve as an intermediate representation that enables consistent modeling of vision-language-action (VLA) models across heterogeneous datasets. However, approaches to…