7 papers · 1 filter
Orbis 2: A Hierarchical World Model for Driving
Sudhanshu Mittal, Arian Mousakhan, Silvio Galesso +4
Current world models operate at a single level of abstraction, with most prioritizing perceptual fidelity while lacking the spatial reasoning and semantic understanding required fo…
The Surprising Effectiveness of Canonical Knowledge Distillation for Semantic Segmentation
Muhammad Ali, Kevin Alexander Laube, Madan Ravi Ganesh +3
Recent knowledge distillation (KD) methods for semantic segmentation introduce increasingly complex hand-crafted objectives, yet are typically evaluated under fixed iteration sched…
What Drives Compositional Generalization? The Importance of Continuous Training Objectives in Visual Generative Models
Karim Farid, Rajat Sahay, Yumna Ali Alnaggar +4
Compositional generalization, the ability to generate novel combinations of known concepts, is a key ingredient for visual generative models. Yet, not all mechanisms that enable or…
Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World Models
Arian Mousakhan, Sudhanshu Mittal, Silvio Galesso +2
Existing world models for autonomous driving struggle with long-horizon generation and generalization to challenging scenarios. In this work, we develop a model using simple design…
Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models
Simon Schrodi, David T. Hoffmann, Max Argus +2
Contrastive vision-language models (VLMs), like CLIP, have gained popularity for their versatile applicability to various downstream tasks. Despite their successes in some tasks, l…
Amodal Optical Flow
Maximilian Luz, Rohit Mohan, Ahmed Rida Sekkat +4
Optical flow estimation is very challenging in situations with transparent or occluded objects. In this work, we address these challenges at the task level by introducing Amodal Op…