1 paper · 1 filter
Zuojin Tang, Feifan Luo, Haoyun Liu +10
Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant observations of physical change. Latent action…