1 paper
Yongjie Bai, Hanting Wang, Mingtong Dai +3
General-purpose vision-language-action models benefit from large vision-language priors, but effective manipulation also requires anticipating action-relevant scene changes. Existi…