12 papers
VOLA: Improving Open-World Driving by VLM-Based Semantic Attribute Prediction
Yuchen Zhang, Yuan Gao, Sebastian Schmidt +1
Driving in the real world is open-world: a car may encounter a fallen mattress, a deer, or other objects outside its training data. Naming them is not enough. The system must know…
Imagined Rollouts are Kinematic, Not Dynamic: A Diagnosis of Long-Horizon World-Model Failure
Finn Rasmus Schäfer, Korbinian Moller, Yuan Gao +3
Long-horizon failure in world models is conventionally attributed to compounding error, a generic framing that does not distinguish what kind of error compounds. We propose a kinem…
EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving
Finn Rasmus Schäfer, Yuan Gao, Dingrui Wang +5
While Vision-Language Models (VLMs) have advanced high-level reasoning in autonomous driving, their ability to ground this reasoning in the underlying physics of ego-motion remains…
Smooth Piecewise Cutting for Neural Operator to Handle Discontinuities and Sharp Transitions
Ha Dang, Sebastian Schmidt, Juergen Hesser
Neural operators have achieved strong performance in learning solution operators of partial differential equations (PDEs), but their inherently continuous representations struggle…
Scalable Object Detection in the Car Interior With Vision Foundation Models
Sebastian Schmidt, Bálint Mészáros, Ahmet Firintepe +1
AI tasks in the car interior like identifying and localizing externally introduced objects is crucial for response quality of personal assistants. However, computational resources…
Amplified Patch-Level Differential Privacy for Free via Random Cropping
Kaan Durmaz, Jan Schuchardt, Sebastian Schmidt +1
Random cropping is one of the most common data augmentation techniques in computer vision, yet the role of its inherent randomness in training differentially private machine learni…