collaborators

6 papers

cs.CV2026

Walking in the Implicit: Interactive World Exploration via Neural Scene Representation

Zhiqi Li, Chengrui Dong, Zhenhua Du +6

Interactive video generation systems for camera-controlled world exploration roll out growing sequences of latent video frames, entangling state transition with high-frequency obse…

cs.RO2026

AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving

Yanhao Wu, Haoyang Zhang, Fei He +6

Practical autonomous driving requires models that generalize by reasoning through spatial-temporal possibilities to exclude unsafe outcomes. While state-of-the-art (SOTA) methods u…

cs.CV2026

UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register

Congpei Qiu, Zhaoyu Hu, Wei Ke +3

Representation learning with Vision Transformers (ViTs) has advanced rapidly, yet the utility of large-scale models in spatially sensitive tasks is hindered by spurious tokens. Pri…

cs.CV2026

Any 3D Scene is Worth 1K Tokens: 3D-Grounded Representation for Scene Generation at Scale

Dongxu Wei, Qi Xu, Zhiqi Li +6

3D scene generation has long been dominated by 2D multi-view or video diffusion models. This is due not only to the lack of scene-level 3D latent representation, but also to the fa…

cs.CV2025

Refining CLIP's Spatial Awareness: A Visual-Centric Perspective

Congpei Qiu, Yanhao Wu, Wei Ke +2

Contrastive Language-Image Pre-training (CLIP) excels in global alignment with language but exhibits limited sensitivity to spatial information, leading to strong performance in ze…

cs.CV2025

Generating Multimodal Driving Scenes via Next-Scene Prediction

Yanhao Wu, Haoyang Zhang, Tianwei Lin +6

Generative models in Autonomous Driving (AD) enable diverse scene creation, yet existing methods fall short by only capturing a limited range of modalities, restricting the capabil…