Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Bridging Structure and Language: Graph-Based Visual Reasoning for Autonomous Road Understanding
Lena Wild, Katie Z Luo, Marco Pavone
Structured road understanding of lane geometry, topology, and traffic element relationships is foundational to safe autonomous driving. While vision-language models (VLMs) offer pr…
cs.CV2025
Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models
Katie Luo, Jingwei Ji, Tong He +4
Current autonomous driving systems rely on specialized models for perceiving and predicting motion, which demonstrate reliable performance in standard conditions. However, generali…
cs.CV2025
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation
Yichen Xie, Runsheng Xu, Tong He +9
The latest advancements in multi-modal large language models (MLLMs) have spurred a strong renewed interest in end-to-end motion planning approaches for autonomous driving. Many en…