activity
20242026
collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2026

XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments

Kangan Qian, ChuChu Xie, Yang Zhong +13

Vision-Language-Action (VLA) models drive next-generation autonomous systems, but training them requires scalable, high-quality annotations from complex environments. Current cloud…

cs.CV2025

Realistic and Controllable 3D Gaussian-Guided Object Editing for Driving Video Generation

Jiusi Li, Jackson Jiang, Jinyu Miao +10

Corner cases are crucial for training and validating autonomous driving systems, yet collecting them from the real world is often costly and hazardous. Editing objects within captu…

cs.CV2025

How Cars Move: Analyzing Driving Dynamics for Safer Urban Traffic

Kangan Qian, Jinyu Miao, Xinyu Jiao +6

Understanding the spatial dynamics of cars within urban systems is essential for optimizing infrastructure management and resource allocation. Recent empirical approaches for analy…

cs.CV2025

A Survey on Vision-Language-Action Models for Autonomous Driving

Sicong Jiang, Zilin Huang, Kangan Qian +17

The rapid progress of multimodal large language models (MLLM) has paved the way for Vision-Language-Action (VLA) paradigms, which integrate visual perception, natural language unde…

cs.CV2025

COME: Adding Scene-Centric Forecasting Control to Occupancy World Model

Yining Shi, Kun Jiang, Qiang Meng +6

World models are critical for autonomous driving to simulate environmental dynamics and generate synthetic data. Existing methods struggle to disentangle ego-vehicle motion (perspe…

cs.CV2025

CleanMAP: Distilling Multimodal LLMs for Confidence-Driven Crowdsourced HD Map Updates

Ankit Kumar Shaw, Kun Jiang, Tuopu Wen +5

The rapid growth of intelligent connected vehicles (ICVs) and integrated vehicle-road-cloud systems has increased the demand for accurate, real-time HD map updates. However, ensuri…