1 paper
Wanshun Xu, Long Zhuang, Lianlei Shan
Vision-Language-Action (VLA) models offer a unified framework for robotic perception and control, but their ability to scale to real-world, long-horizon tasks is limited by the hig…