1 paper
Yihao Wang, Pengxiang Ding, Lingxiao Li +13
Vision-Language-Action (VLA) models typically bridge the gap between perceptual and action spaces by pre-training a large-scale Vision-Language Model (VLM) on robotic data. While t…