3 papers
cs.RO2026
GeoAlign: Beyond Semantics with State-Guided Spatial Alignment in VLA Models
Yizhi Chen, Zhanxiang Cao, Xinyi Peng +14
Current Vision--Language--Action (VLA) models often optimize for semantic grounding, whereas executable manipulation requires geometry-aware spatial alignment and dynamic affordanc…
cs.RO2026
Driving Intents Amplify Planning-Oriented Reinforcement Learning
Hengtong Lu, Victor Shea-Jay Huang, Chengmin Yang +4
Continuous-action policies trained on a single demonstrated trajectory per scene suffer from mode collapse: samples cluster around the demonstrated maneuver and the policy cannot r…
cs.RO2026
MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous Driving
Yuzhou Huang, Benjin Zhu, Hengtong Lu +6
Autonomous driving has progressed from modular pipelines toward end-to-end unification, and Vision-Language-Action (VLA) models are a natural extension of this journey beyond Visio…