3 papers
cs.RO2026
Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking
Jinquan Zhang, Dongfu Yin, Run Yang +3
Vision-Language-Action (VLA) policies promise general robotic manipulation, but their robustness against physical-world attacks remains fragile. In particular, we show that physica…
cs.AI2026
Beyond Pixels: Vector-to-Graph Transformation for Reliable Schematic Auditing
Chengwei Ma, Zhen Tian, Zhou Zhou +5
Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual understanding, yet they suffer from a critical limitation: structural blindness. Even state-of-the…
cs.CV2026
AugVLA-3D: Depth-Driven Feature Augmentation for Vision-Language-Action Models
Zhifeng Rao, Wenlong Chen, Lei Xie +4
Vision-Language-Action (VLA) models have recently achieved remarkable progress in robotic perception and control, yet most existing approaches primarily rely on VLM trained using 2…