VLATest: Testing and Evaluating Vision-Language-Action Models for Robotic Manipulation
arXiv:2409.12894 · doi:10.1145/3729343
Abstract
The rapid advancement of generative AI and multi-modal foundation models has shown significant potential in advancing robotic manipulation. Vision-language-action (VLA) models, in particular, have emerged as a promising approach for visuomotor control by leveraging large-scale vision-language data and robot demonstrations. However, current VLA models are typically evaluated using a limited set of hand-crafted scenes, leaving their general performance and robustness in diverse scenarios largely unexplored. To address this gap, we present VLATest, a fuzzing framework designed to generate robotic manipulation scenes for testing VLA models. Based on VLATest, we conducted an empirical study to assess the performance of seven representative VLA models. Our study results revealed that current VLA models lack the robustness necessary for practical deployment. Additionally, we investigated the impact of various factors, including the number of confounding objects, lighting conditions, camera poses, unseen objects, and task instruction mutations, on the VLA model's performance. Our findings highlight the limitations of existing VLA models, emphasizing the need for further research to develop reliable and trustworthy VLA applications.
To appear in FSE '25 (Proceedings of ACM Software Engineering, Vol. 2, Issue FSE, Article FSE073), 24 pages, 7 figures
References in corpus (10)
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- DINOv2: Learning Robust Visual Features without Supervision
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- The GOOSE Dataset for Perception in Unstructured Environments
- Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models
- Testing of Deep Reinforcement Learning Agents with Surrogate Models
- DeepGD: A Multi-Objective Black-Box Test Selection Approach for Deep Neural Networks
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- FaithDial: A Faithful Benchmark for Information-Seeking Dialogue
- MultiPL-E: A Scalable and Extensible Approach to Benchmarking Neural Code Generation