21 citations · 24 across the 3 of their papers we have counts for
5 papers
On the Road with GPT-4V(ision): Early Explorations of Visual-Language Model on Autonomous Driving
Licheng Wen, Xuemeng Yang, Daocheng Fu +15
The pursuit of autonomous driving technology hinges on the sophisticated integration of perception, decision-making, and control systems. Traditional approaches, both data-driven a…
SceneDM: Scene-level Multi-agent Trajectory Generation with Consistent Diffusion Models
Zhiming Guo, Xing Gao, Jianlan Zhou +2
Realistic scene-level multi-agent motion simulations are crucial for developing and evaluating self-driving algorithms. However, most existing works focus on generating trajectorie…
Homogeneous Multi-modal Feature Fusion and Interaction for 3D Object Detection
Xin Li, Botian Shi, Yuenan Hou +4
Multi-modal 3D object detection has been an active research topic in autonomous driving. Nevertheless, it is non-trivial to explore the cross-modal feature fusion between sparse 3D…
A Benchmark for Structured Procedural Knowledge Extraction from Cooking Videos
Frank F. Xu, Lei Ji, Botian Shi +4
Watching instructional videos are often used to learn about procedures. Video captioning is one way of automatically collecting such knowledge. However, it provides only an indirec…
UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation
Huaishao Luo, Lei Ji, Botian Shi +6
With the recent success of the pre-training technique for NLP and image-linguistic tasks, some video-linguistic pre-training works are gradually developed to improve video-text rel…