2 papers
cs.CV2026
A Collaborative Multi-Modality Interaction for VLA-based End-to-End Autonomous Driving
Jingtao Sun, Xiaohai He, Yike Zhang +4
Vision-Language-Action (VLA) models have emerged as a powerful paradigm for end-to-end autonomous driving by jointly integrating perception, reasoning, and decision making within a…
cs.RO2026
CRAFT: Adapting VLA Models to Contact-rich Manipulation via Force-aware Curriculum Fine-tuning
Yike Zhang, Yaonan Wang, Xinxin Sun +6
Vision-Language-Action (VLA) models have shown a strong capability in enabling robots to execute general instructions, yet they struggle with contact-rich manipulation tasks, where…