7 papers
Scaling Self-Play for End-to-End Driving
Luke Rowe, Roger Girgis, Rodrigue de Schaetzen +6
End-to-end autonomous driving models are typically trained on offline human-demonstration datasets that provide limited state coverage and often no closed-loop feedback, making the…
Weak-to-Strong Knowledge Distillation Accelerates Visual Learning
Baiang Li, Wenhao Chai, Felix Heide
Large-scale visual learning is increasingly limited by training cost. Existing knowledge distillation methods transfer from a stronger teacher to a weaker student for compression o…
ScenarioControl: Vision-Language Controllable Vectorized Latent Scenario Generation
Lili Gao, Yanbo Xu, William Koch +8
We introduce ScenarioControl, the first vision-language control mechanism for learned driving scenario generation. Given a text prompt or an input image, Scenario-Control synthesiz…
VERDI: VLM-Embedded Reasoning for Autonomous Driving
Bowen Feng, Zhiting Mei, Julian Ost +5
While autonomous driving (AD) stacks struggle with decision making under partial observability and real-world complexity, human drivers are capable of applying commonsense reasonin…
WorldFlow3D: Flowing Through 3D Distributions for Unbounded World Generation
Amogh Joshi, Julian Ost, Felix Heide
Unbounded 3D world generation is emerging as a foundational task for scene modeling in computer vision, graphics, and robotics. In this work, we present WorldFlow3D, a novel method…
HEIR: Learning Graph-Based Motion Hierarchies
Cheng Zheng, William Koch, Baiang Li +1
Hierarchical structures of motion exist across research fields, including computer vision, graphics, and robotics, where complex dynamics typically arise from coordinated interacti…