most citedA Survey on Vision-Language-Action Models for Autonomous Driving

2 citations · 2 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2025

Realistic and Controllable 3D Gaussian-Guided Object Editing for Driving Video Generation

Jiusi Li, Jackson Jiang, Jinyu Miao +10

Corner cases are crucial for training and validating autonomous driving systems, yet collecting them from the real world is often costly and hazardous. Editing objects within captu…

cs.CV20252 cited

A Survey on Vision-Language-Action Models for Autonomous Driving

Sicong Jiang, Zilin Huang, Kangan Qian +17

The rapid progress of multimodal large language models (MLLM) has paved the way for Vision-Language-Action (VLA) paradigms, which integrate visual perception, natural language unde…

cs.CV2025

COME: Adding Scene-Centric Forecasting Control to Occupancy World Model

Yining Shi, Kun Jiang, Qiang Meng +6

World models are critical for autonomous driving to simulate environmental dynamics and generate synthetic data. Existing methods struggle to disentangle ego-vehicle motion (perspe…

cs.CV2025

CleanMAP: Distilling Multimodal LLMs for Confidence-Driven Crowdsourced HD Map Updates

Ankit Kumar Shaw, Kun Jiang, Tuopu Wen +5

The rapid growth of intelligent connected vehicles (ICVs) and integrated vehicle-road-cloud systems has increased the demand for accurate, real-time HD map updates. However, ensuri…

cs.CV2025

POD: Predictive Object Detection with Single-Frame FMCW LiDAR Point Cloud

Yining Shi, Kun Jiang, Xin Zhao +5

LiDAR-based 3D object detection is a fundamental task in the field of autonomous driving. This paper explores the unique advantage of Frequency Modulated Continuous Wave (FMCW) LiD…

cs.CV2025

LEGO-Motion: Learning-Enhanced Grids with Occupancy Instance Modeling for Class-Agnostic Motion Prediction

Kangan Qian, Jinyu Miao, Ziang Luo +7

Accurate and reliable spatial and motion information plays a pivotal role in autonomous driving systems. However, object-level perception models struggle with handling open scenari…