1 paper
Xiang Zhu, Puzhen Yuan, Yichen Liu +1
Learning generalizable vision-language-action (VLA) models from large-scale human videos is promising but challenging due to cross-embodiment discrepancies in both visual observati…