8 papers
Decompose, Compare, and Decide: Multimodal LLMs are Implicit Few-Shot Learners
Yunhan Wang, Eshika Khandelwal, Edson Araujo +3
Multimodal Large Language Models (MLLMs) have demonstrated remarkable abilities when analyzing images, yet translating these capabilities to few-shot image classification remains c…
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
Walid Bousselham, Hilde Kuehne, Cordelia Schmid
Training vision-language models (VLMs) for complex reasoning remains a challenging task, i.a. due to the scarcity of high-quality image-text reasoning data. Conversely, text-based…
DEX-AR: A Dynamic Explainability Method for Autoregressive Vision-Language Models
Walid Bousselham, Angie Boggust, Hendrik Strobelt +1
As Vision-Language Models (VLMs) become increasingly sophisticated and widely used, it becomes more and more crucial to understand their decision-making process. Traditional explai…
MaskInversion: Localized Embeddings via Optimization of Explainability Maps
Walid Bousselham, Sofian Chaybouti, Christian Rupprecht +2
Vision-language foundation models such as CLIP have achieved tremendous results in global vision-language alignment, but still show some limitations in creating representations for…
When LLaVA Meets Objects: Token Composition for Vision-Language-Models
Soumya Jahagirdar, Walid Bousselham, Anna Kukleva +1
Current autoregressive Vision Language Models (VLMs) usually rely on a large number of visual tokens to represent images, resulting in a need for more compute especially at inferen…
REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
Sofian Chaybouti, Walid Bousselham, Moritz Wolter +1
Video-Question-Answering (VideoQA) comprises the capturing of complex visual relation changes over time, remaining a challenge even for advanced Video Language Models (VLM), i.a.,…