3 papers
cs.CV2025
VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision
Yi Xu, Yuxin Hu, Zaiwei Zhang +5
Human drivers rely on commonsense reasoning to navigate diverse and dynamic real-world scenarios. Existing end-to-end (E2E) autonomous driving (AD) models are typically optimized t…
cs.CV2024
Flash3D: Super-scaling Point Transformers through Joint Hardware-Geometry Locality
Liyan Chen, Gregory P. Meyer, Zaiwei Zhang +2
Recent efforts recognize the power of scale in 3D learning (e.g. PTv3) and attention mechanisms (e.g. FlashAttention). However, current point cloud backbones fail to holistically u…
cs.CV2024
Uncertainty-Guided Enhancement on Driving Perception System via Foundation Models
Yunhao Yang, Yuxin Hu, Mao Ye +5
Multimodal foundation models offer promising advancements for enhancing driving perception systems, but their high computational and financial costs pose challenges. We develop a m…