4 papers
DriveTok: 3D Driving Scene Tokenization for Unified Multi-View Reconstruction and Understanding
Dong Zhuo, Wenzhao Zheng, Sicheng Zuo +4
With the growing adoption of vision-language-action models and world models in autonomous driving systems, scalable image tokenization becomes crucial as the interface for the visu…
ViGoR: Improving Visual Grounding of Large Vision Language Models with Fine-Grained Reward Modeling
Siming Yan, Min Bai, Weifeng Chen +3
By combining natural language understanding, generation capabilities, and breadth of knowledge of large language models with image perception, recent large vision language models (…
DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning
Zhe Liu, Runhui Huang, Rui Yang +6
Although multi-modal large language models (MLLMs) have shown strong capabilities across diverse domains, their application in generating fine-grained 3D perception and prediction…
Wavelet-based Decoupling Framework for low-light Stereo Image Enhancement
Shuangli Du, Siming Yan, Zhenghao Shi +2
Low-light images suffer from complex degradation, and existing enhancement methods often encode all degradation factors within a single latent space. This leads to highly entangled…