1 paper
Bin Yu, Yao Zhang, Haishan Liu +9
Vision-language-action (VLA) models are built on the premise that semantic understanding from pretrained language or vision-language backbones should guide robot action prediction.…