8 papers
EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving
Finn Rasmus Schäfer, Yuan Gao, Dingrui Wang +5
While Vision-Language Models (VLMs) have advanced high-level reasoning in autonomous driving, their ability to ground this reasoning in the underlying physics of ego-motion remains…
WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching
Weilun Feng, Guoxin Fan, Haotong Qin +10
Diffusion-based world models have shown strong potential for unified world simulation, but the iterative denoising remains too costly for interactive use and long-horizon rollouts.…
Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets?
Dingrui Wang, Zhihao Liang, Hongyuan Ye +13
While recent video world models can generate highly realistic videos, their ability to perform semantic reasoning and planning remains unclear and unquantified. We introduce Target…
DUPLEX: Agentic Dual-System Planning via LLM-Driven Information Extraction
Keru Hua, Ding Wang, Yaoying Gu +1
While Large Language Models (LLMs) provide semantic flexibility for robotic task planning, their susceptibility to hallucination and logical inconsistency limits their reliability…
Foundation Models in Autonomous Driving: A Survey on Scenario Generation and Scenario Analysis
Yuan Gao, Mattia Piccinini, Yuchen Zhang +12
For autonomous vehicles, safe navigation in complex environments depends on handling a broad range of diverse and rare driving scenarios. Simulation- and scenario-based testing hav…
Enhancing Physical Consistency in Lightweight World Models
Dingrui Wang, Zhexiao Sun, Zhouheng Li +8
A major challenge in deploying world models is the trade-off between size and performance. Large world models can capture rich physical dynamics but require massive computing resou…