activity
20212026
most citedDeepInteraction: 3D Object Detection via Modality Interaction

65 citations · 220 across the 34 of their papers we have counts for

collaborators

34 papers

cs.CV2026

Metis: A Generalizable and Efficient World-Action Model for Autonomous Driving and Urban Navigation

Jingyu Li, Zhe Liu, Dongnan Hu +10

World action models~(WAMs) have shown great promise for autonomous driving and urban navigation. Built upon Vision-Language-Action models or video generation models, existing appro…

cs.RO2026

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation

Yajie Li, Bozhou Zhang, Chun Gu +5

Video generation models offer a promising imagination mechanism for robot manipulation by predicting long-horizon future observations, but effectively exploiting these imagined fut…

cs.CV2026

See Tomorrow, Act Today: Foresight-Driven Autonomous Driving

Bozhou Zhang, Nan Song, Yuang Wang +3

Current end-to-end autonomous driving planners are fundamentally reactive: they condition on historical and present observations to predict future actions. We argue that autonomous…

cs.RO2026

Uni-World VLA: Interleaved World Modeling and Planning for Autonomous Driving

Qiqi Liu, Huan Xu, Jingyu Li +5

Autonomous driving requires reasoning about how the environment evolves and planning actions accordingly. Existing world-model-based approaches typically predict future scenes firs…

cs.RO2026

UniMotion: A Unified Motion Framework for Simulation, Prediction and Planning

Nan Song, Junzhe Jiang, Jingyu Li +2

Motion simulation, prediction and planning are foundational tasks in autonomous driving, each essential for modeling and reasoning about dynamic traffic scenarios. While often addr…

cs.CV2026

SGDrive: Scene-to-Goal Hierarchical World Cognition for Autonomous Driving

Jingyu Li, Junjie Wu, Dongnan Hu +6

Recent end-to-end autonomous driving approaches have leveraged Vision-Language Models (VLMs) to enhance planning capabilities in complex driving scenarios. However, VLMs are inhere…