2 citations · 2 across the 8 of their papers we have counts for
16 papers · 1 filter
DriveJudge: Rethinking Autonomous Driving Evaluation with Vision-Language Models
Xinglong Sun, Kevin Xie, Jenny Schmalfuss +5
Autonomous driving has shifted towards end-to-end policy learning, where reliable, interpretable policy evaluation is a fundamental challenge as driving quality is highly context-d…
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
Shihao Wang, Guo Chen, De-an Huang +6
While Video Large Language Models (Video-LLMs) have shown significant potential in multimodal understanding and reasoning tasks, how to efficiently select the most informative fram…
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
Shihao Wang, Zhiding Yu, Xiaohui Jiang +6
The advances in vision-language models (VLMs) have led to a growing interest in autonomous driving to leverage their strong reasoning capabilities. However, extending these capabil…
Hydra-MDP++: Advancing End-to-End Driving via Expert-Guided Hydra-Distillation
Kailin Li, Zhenxin Li, Shiyi Lan +6
Hydra-MDP++ introduces a novel teacher-student knowledge distillation framework with a multi-head decoder that learns from human demonstrations and rule-based experts. Using a ligh…
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
Zhiqi Li, Guo Chen, Shilong Liu +24
Recently, promising progress has been made by open-source vision-language models (VLMs) in bringing their capabilities closer to those of proprietary frontier models. However, most…
Exploring Camera Encoder Designs for Autonomous Driving Perception
Barath Lakshmanan, Joshua Chen, Shiyi Lan +3
The cornerstone of autonomous vehicles (AV) is a solid perception system, where camera encoders play a crucial role. Existing works usually leverage pre-trained Convolutional Neura…