activity
20202023
most citedOn the Road with GPT-4V(ision): Early Explorations of Visual-Language Model on Autonomous Driving

21 citations · 24 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CV202321 cited

On the Road with GPT-4V(ision): Early Explorations of Visual-Language Model on Autonomous Driving

Licheng Wen, Xuemeng Yang, Daocheng Fu +15

The pursuit of autonomous driving technology hinges on the sophisticated integration of perception, decision-making, and control systems. Traditional approaches, both data-driven a…

cs.RO20231 cited

SceneDM: Scene-level Multi-agent Trajectory Generation with Consistent Diffusion Models

Zhiming Guo, Xing Gao, Jianlan Zhou +2

Realistic scene-level multi-agent motion simulations are crucial for developing and evaluating self-driving algorithms. However, most existing works focus on generating trajectorie…

cs.CV20222 cited

Homogeneous Multi-modal Feature Fusion and Interaction for 3D Object Detection

Xin Li, Botian Shi, Yuenan Hou +4

Multi-modal 3D object detection has been an active research topic in autonomous driving. Nevertheless, it is non-trivial to explore the cross-modal feature fusion between sparse 3D…

cs.CL2020

A Benchmark for Structured Procedural Knowledge Extraction from Cooking Videos

Frank F. Xu, Lei Ji, Botian Shi +4

Watching instructional videos are often used to learn about procedures. Video captioning is one way of automatically collecting such knowledge. However, it provides only an indirec…

cs.CV2020

UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

Huaishao Luo, Lei Ji, Botian Shi +6

With the recent success of the pre-training technique for NLP and image-linguistic tasks, some video-linguistic pre-training works are gradually developed to improve video-text rel…