1 paper
Teli Ma, Jia Zheng, Zifan Wang +4
Vision-Language-Action (VLA) models have emerged as a promising paradigm for robot learning, but their representations are still largely inherited from static image-text pretrainin…