3 papers
cs.CV2026
ProSR: Process-Shaped Spatial Reasoning for Reliable Chain-of-Thought in VLMs
Jiangyang Li, Cong Wan, Changjie Wu +8
Reliable spatial reasoning remains a core bottleneck for vision-language models (VLMs). Existing mainstream training paradigms for spatial reasoning largely rely on outcome alignme…
cs.CV2026
DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving
Lingjun Zhang, Changjie Wu, Linzhe Shi +6
End-to-end autonomous driving systems are increasingly integrating Vision-Language Model (VLM) architectures, incorporating text reasoning or visual reasoning to enhance the robust…
cs.CV2026
MindDriver: Introducing Progressive Multimodal Reasoning for Autonomous Driving
Lingjun Zhang, Yujian Yuan, Changjie Wu +7
Vision-Language Models (VLM) exhibit strong reasoning capabilities, showing promise for end-to-end autonomous driving systems. Chain-of-Thought (CoT), as VLM's widely used reasonin…