6 papers
SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models
Olaf Dünkel, Basavaraj Sunagad, Haoran Wang +3
Measuring structured object understanding in vision foundation models remains challenging due to inconsistent evaluation protocols and limited part-level supervision. Semantic corr…
No Hard Negatives Required: Concept Centric Learning Leads to Compositionality without Degrading Zero-shot Capabilities of Contrastive Models
Hai X. Pham, David T. Hoffmann, Ricardo Guerrero +1
Contrastive vision-language (V&L) models remain a popular choice for various applications. However, several limitations have emerged, most notably the limited ability of V&L models…
Unlocking In-Context Learning for Natural Datasets Beyond Language Modelling
Jelena BratuliÄ, Sudhanshu Mittal, David T. Hoffmann +5
Large Language Models (LLMs) exhibit In-Context Learning (ICL), which enables the model to perform new tasks conditioning only on the examples provided in the context without updat…
Common Data Properties Limit Object-Attribute Binding in CLIP
Bijay Gurung, David T. Hoffmann, Thomas Brox
Contrastive vision-language models like CLIP are used for a large variety of applications, such as zero-shot classification or as vision encoder for multi-modal models. Despite the…
Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models
Simon Schrodi, David T. Hoffmann, Max Argus +2
Contrastive vision-language models (VLMs), like CLIP, have gained popularity for their versatile applicability to various downstream tasks. Despite their successes in some tasks, l…
Floxels: Fast Unsupervised Voxel Based Scene Flow Estimation
David T. Hoffmann, Syed Haseeb Raza, Hanqiu Jiang +3
Scene flow estimation is a foundational task for many robotic applications, including robust dynamic object detection, automatic labeling, and sensor synchronization. Two types of…