5 papers
Principles of Robot Autonomy
Daniele Gammelli, Joseph Lorenzetti, Katie Luo +2
Autonomous robots are moving rapidly from research labs into everyday life - on roads, in the air, in warehouses, and in space. Robot autonomy is no longer solely an academic pursu…
DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners?
Jadelynn Dao, Milan Ganai, Yasmina Abukhadra +7
Vision-Language Models (VLMs) are increasingly deployed as high-level planners for embodied agents, with an emerging strategy of scaling test-time compute to improve capability. Ho…
Bridging Structure and Language: Graph-Based Visual Reasoning for Autonomous Road Understanding
Lena Wild, Katie Z Luo, Marco Pavone
Structured road understanding of lane geometry, topology, and traffic element relationships is foundational to safe autonomous driving. While vision-language models (VLMs) offer pr…
Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models
Katie Luo, Jingwei Ji, Tong He +4
Current autonomous driving systems rely on specialized models for perceiving and predicting motion, which demonstrate reliable performance in standard conditions. However, generali…
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation
Yichen Xie, Runsheng Xu, Tong He +9
The latest advancements in multi-modal large language models (MLLMs) have spurred a strong renewed interest in end-to-end motion planning approaches for autonomous driving. Many en…