6 papers
ProSR: Process-Shaped Spatial Reasoning for Reliable Chain-of-Thought in VLMs
Jiangyang Li, Cong Wan, Changjie Wu +8
Reliable spatial reasoning remains a core bottleneck for vision-language models (VLMs). Existing mainstream training paradigms for spatial reasoning largely rely on outcome alignme…
ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models
Zuojin Tang, Haoyun Liu, Xinyuan Chang +11
Vision-language-action (VLA) models remain constrained by the scarcity of action-labeled robot data, whereas action-free videos provide abundant evidence of how the physical world…
DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving
Lingjun Zhang, Changjie Wu, Linzhe Shi +6
End-to-end autonomous driving systems are increasingly integrating Vision-Language Model (VLM) architectures, incorporating text reasoning or visual reasoning to enhance the robust…
Persistent Autoregressive Mapping with Traffic Rules for Autonomous Driving
Shiyi Liang, Xinyuan Chang, Changjie Wu +8
Safe autonomous driving requires both accurate HD map construction and persistent awareness of traffic rules, even when their associated signs are no longer visible. However, exist…
MindDriver: Introducing Progressive Multimodal Reasoning for Autonomous Driving
Lingjun Zhang, Yujian Yuan, Changjie Wu +7
Vision-Language Models (VLM) exhibit strong reasoning capabilities, showing promise for end-to-end autonomous driving systems. Chain-of-Thought (CoT), as VLM's widely used reasonin…
UniMapGen: A Generative Framework for Large-Scale Map Construction from Multi-modal Data
Yujian Yuan, Changjie Wu, Xinyuan Chang +6
Large-scale map construction plays a vital role in applications like autonomous driving and navigation systems. Traditional large-scale map construction approaches mainly rely on c…