Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
Why do LLaVA Vision-Language Models Reply to Images in English?
Musashi Hinck, Carolin Holtermann, Matthew Lyle Olson +6
We uncover a surprising multilingual bias occurring in a popular class of multimodal vision-language models (VLMs). Including an image in the query to a LLaVA-style VLM significant…
cs.CL2024
What the Weight?! A Unified Framework for Zero-Shot Knowledge Composition
Carolin Holtermann, Markus Frohmann, Navid Rekabsaz +1
The knowledge encapsulated in a model is the core factor determining its final performance on downstream tasks. Much research in NLP has focused on efficient methods for storing an…