3 papers
cs.CV2026
DriveStack-VLA: Render-Teacher Alignment for BEV-Based DeepStack Vision-Language-Action Model
Jingke Wang, Zhenru Zhao, Shuangming Lei +8
Vision-Language-Action driving models convert a pretrained Vision-Language Model into a driving policy, allowing them to use world knowledge and follow language guidances. However,…
cs.CV2026
SparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene Representation
Ruoyu Wang, Jingke Wang, Yukai Ma +5
Recently, world models have made significant progress in enhancing end-to-end driving systems through both future situation forecasting and improved scene understanding. However, e…
cs.CV2026
SSR: Pushing the Limit of Spatial Intelligence with Structured Scene Reasoning
Yi Zhang, Youya Xia, Yong Wang +7
While Multimodal Large Language Models (MLLMs) excel in semantic tasks, they frequently lack the "spatial sense" essential for sophisticated geometric reasoning. Current models typ…