7 papers
4D-WAM: 4D Consistent World Modeling for Autonomous Driving
Jiacheng Fu, Yibo Yuan, Meng Tian +8
Emerging World-Action Models (WAMs) have demonstrated promising performance in autonomous driving by jointly modeling future driving scene evolution and trajectory planning. Howeve…
SUV: Future Scene Understanding as Video Generation for End-to-End Driving
Yibo Yuan, Jiacheng Fu, Jiangtong Zhu +8
End-to-end driving requires a coherent understanding of future scenes, yet existing methods model these scenes using task-specific heads and output formats, with limited scalabilit…
Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search
Xinlei Yin, Xiulian Peng, Xiao Li +2
Long video understanding presents significant challenges for vision-language models due to extremely long context windows. Existing solutions relying on naive chunking strategies w…
EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models
Wenhao Xu, Xin Dong, Yue Li +2
Video large language models have demonstrated strong video understanding capabilities but suffer from high inference costs due to the massive number of tokens in long videos. Inspi…
3D Reconstruction from Transient Measurements with Time-Resolved Transformer
Yue Li, Shida Sun, Yu Hong +2
Transient measurements, captured by the timeresolved systems, are widely employed in photon-efficient reconstruction tasks, including line-of-sight (LOS) and non-line-of-sight (NLO…
Drive-R1: Bridging Reasoning and Planning in VLMs for Autonomous Driving with Reinforcement Learning
Yue Li, Meng Tian, Dechang Zhu +4
Large vision-language models (VLMs) for autonomous driving (AD) are evolving beyond perception and cognition tasks toward motion planning. However, we identify two critical challen…