2 papers
cs.RO2025
An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
Chao Xu, Suyu Zhang, Yang Liu +11
Vision-Language-Action (VLA) models are driving a revolution in robotics, enabling machines to understand instructions and interact with the physical world. This field is exploding…
cs.CV2025
Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents
Dayong Liu, Chao Xu, Weihong Chen +5
Multimodal Large Language Models (MLLMs) show promising results as decision-making engines for embodied agents operating in complex, physical environments. However, existing benchm…