4 papers
SafeVLA-Bench: A Benchmark for the Success-Safety Gap in Vision-Language-Action Models
Jialiang Fan, Weizhe Xu, Oleg Sokolsky +2
Vision-language-action (VLA) benchmarks measure whether a policy completes a requested manipulation task, but binary success can hide safety-relevant trajectory behavior: reaching…
SafePilot: A Framework for Assuring LLM-enabled Cyber-Physical Systems
Weizhe Xu, Mengyu Liu, Fanxin Kong
Large Language Models (LLMs), deep learning architectures with typically over 10 billion parameters, have recently begun to be integrated into various cyber-physical systems (CPS)…
SafeGen-LLM: Enhancing Safety Generalization in Task Planning for Robotic Systems
Jialiang Fan, Weizhe Xu, Mengyu Liu +3
Safety-critical task planning in robotic systems remains challenging: classical planners suffer from poor scalability, Reinforcement Learning (RL)-based methods generalize poorly,…
Enhancing LLM-Based Test Generation by Eliminating Covered Code
WeiZhe Xu, Mengyu Liu, Fanxin Kong
Automated test generation is essential for software quality assurance, with coverage rate serving as a key metric to ensure thorough testing. Recent advancements in Large Language…