7 papers
X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving
Chaoda Zheng, Sean Li, Jinhao Deng +9
Scalable and reliable evaluation is increasingly critical in the end-to-end era of autonomous driving, where vision--language--action (VLA) policies directly map raw sensor streams…
FutureX: Enhance End-to-End Autonomous Driving via Latent Chain-of-Thought World Model
Hongbin Lin, Yiming Yang, Yifan Zhang +10
In autonomous driving, end-to-end planners learn scene representations from raw sensor data and utilize them to generate a motion plan or control actions. However, exclusive relian…
DriveFlow: Rectified Flow Adaptation for Robust 3D Object Detection in Autonomous Driving
Hongbin Lin, Yiming Yang, Chaoda Zheng +7
In autonomous driving, vision-centric 3D object detection recognizes and localizes 3D objects from RGB images. However, due to high annotation costs and diverse outdoor scenes, tra…
FASTopoWM: Fast-Slow Lane Segment Topology Reasoning with Latent World Models
Yiming Yang, Hongbin Lin, Yueru Luo +7
Lane segment topology reasoning provides comprehensive bird's-eye view (BEV) road scene understanding, which can serve as a key perception module in planning-oriented end-to-end au…
TopoStreamer: Temporal Lane Segment Topology Reasoning in Autonomous Driving
Yiming Yang, Yueru Luo, Bingkun He +8
Lane segment topology reasoning constructs a comprehensive road network by capturing the topological relationships between lane segments and their semantic types. This enables end-…
DriveGEN: Generalized and Robust 3D Detection in Driving via Controllable Text-to-Image Diffusion Generation
Hongbin Lin, Zilu Guo, Yifan Zhang +5
In autonomous driving, vision-centric 3D detection aims to identify 3D objects from images. However, high data collection costs and diverse real-world scenarios limit the scale of…