1 paper · 1 filter
Zhuofan Li, Hongkun Yang, Zhenyang Chen +4
Vision-Language-Action (VLA) models have recently enabled embodied agents to perform increasingly complex tasks by jointly reasoning over visual, linguistic, and motor modalities.…