3 papers
cs.CV2026
Targeted Structure Completion for Sparse-View 3D Reconstruction in Autonomous Driving
Guoqing Wang, Pin Tang, Xiangxuan Ren +2
Reconstructing 3D scene structures from sparse, low-overlap observations remains a fundamental challenge in autonomous driving. Recent state-of-the-art frameworks achieve promising…
cs.CV2026
Grounding Everything in Tokens for Multimodal Large Language Models
Xiangxuan Ren, Zhongdao Wang, Liping Hou +3
Multimodal large language models (MLLMs) have made significant advancements in vision understanding and reasoning. However, the autoregressive Transformer architecture used by MLLM…
cs.CV2025
GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation
Zhenya Yang, Zhe Liu, Yuxiang Lu +6
Physics-aware driving world model is essential for drive planning, out-of-distribution data synthesis, and closed-loop evaluation. However, existing methods often rely on a single…