1 paper
Xinyi Xie, Zican Hu, Zhanyu Liu +7
Vision-Language-Action (VLA) models acquire broad embodied capabilities through large-scale pretraining, yet their generalization remains far more fragile than that of LLMs and VLM…