1 paper
Hengyi Xie, Chenfei Yao, Xianjin Wu +4
Vision-language-action (VLA) models commonly adopt an LLM-centric V→L→A pathway, where visual observations are projected into the representation space of a large language…