1 paper · 1 filter
Yufeng Ji, Wenhao Tang, Haoyi Niu +3
Action supervision in vision-language-action (VLA) models is often treated as a downstream objective for learning action prediction. In this paper, we study it instead as a force t…